Hey Investigator,
Almost every Windows forensic finding passes through the registry at some point: what ran, what persisted, which USB stick was plugged in, which network the machine joined. Yet most investigators meet it through a tool that shows a tree, and never see what is underneath.
This guide takes it apart completely. It starts with what the registry actually is and why Windows cannot function without it, then goes down to the bytes: hive files, the base block, cells, the records that are keys and values, and finally the transaction logs: the part almost every tool ignores, and the reason a hive on disk is often not the registry that was really running.
No prior knowledge assumed. Every byte shown below was read out of a real SYSTEM hive captured from a running machine, so you can follow along in a hex editor.
Where these figures come from. Every number, byte map and table on this page was measured from a live Windows 11 system, its hives and transaction logs captured together through a single Volume Shadow Copy. That system is called the reference system throughout. Identifying values are redacted; every offset, length, count and timestamp is the real measurement.
1. What Is the Registry?
Before any byte-level work, one shared picture: the registry is a hierarchical database of configuration. Programs, drivers and Windows itself store settings in it instead of in their own scattered files.
Think of it as the building's central records office. Every department could keep its own filing cabinet in its own room, and once did, but then nothing can be indexed, secured, or backed up as a unit.
What came before, and why it failed
Early Windows kept configuration in .INI text files: WIN.INI, SYSTEM.INI, and one per application. That approach broke down for reasons worth naming, because the registry is a direct answer to each:
No data types
An INI file stores text. Everything, whether a number, a flag or a binary blob, arrives as a string, and every program has to parse and validate it itself. The registry has typed values: REG_DWORD, REG_SZ, REG_BINARY, REG_MULTI_SZ and others, so the type travels with the data.
No security
An INI file's protection is the file's own permissions: all or nothing. The registry attaches a security descriptor per key, so one branch can be writable by an administrator and readable by everyone, and Windows can enforce it.
No atomicity
Interrupt a program writing an INI file and you get half a file. The registry is transactional: the subject of most of this guide, so an interrupted write can be detected and undone.
No timestamps
An INI file has one modification time for the whole file. The registry stores a last-written timestamp on every key, which is the only per-record time it keeps, and the backbone of registry timelining. Read it carefully though: it records that something about that key changed, not what. See section 9.
The shape of it
The registry is a tree. Keys are the folders; they can contain other keys. Values are the leaves; each has a name, a type and data. A path like HKLM\SYSTEM\CurrentControlSet\Services\Tcpip is a chain of keys, and inside the last one sit values such as Start or ImagePath.
The names you know, such as HKEY_LOCAL_MACHINE and HKEY_CURRENT_USER, are called root keys or hives-in-the-abstract. section 4 shows what they really are: mount points onto files.
2. Why Windows Cannot Function Without It
It is tempting to think of the registry as "settings": the things you change in a preferences dialog. It is far more load-bearing than that. Windows reads the registry before it has a desktop, before it has services, and before most of the operating system exists.
Consider what the kernel needs to know at boot, before anything else can happen:
- Which drivers to load, and in what order. A disk controller driver must load before the disk is usable. That list lives in
SYSTEM\CurrentControlSet\Services, with each driver'sStartvalue giving its load phase (boot, system, auto, manual, disabled). - Which hardware profile and control set are valid.
SYSTEM\Selectnames the current control set, and the LastKnownGood one, which is what "Last Known Good Configuration" recovery actually switches to. - Which services to start, under which accounts. Service configuration, dependencies and logon identities are all registry values.
- How to interpret the filesystem and devices. Mounted device mappings, drive letters and volume identities are registry data.
The consequence: the SYSTEM hive is loaded by the boot loader itself and handed to the kernel in memory. It is not a file the operating system opens once it is running . It is one of the things that makes it run. A corrupt SYSTEM hive is not a misconfigured machine; it is a machine that does not boot.
This is also why the registry needs the crash-safety machinery in sections 15 to 17. A configuration file that can be half-written is an annoyance. A boot-path database that can be half-written is an unbootable computer.
3. Why It Matters in Digital Forensics
Logs record what a system chose to write down. The registry records what a system was configured to do, and it keeps that record whether or not anyone intended it as evidence.
Execution
UserAssist counts and times GUI program launches. BAM and DAM record executable paths with last-run times per user SID. Prefetch's configuration, ShimCache and AmCache all live in or beside the registry. Explained in full: AmCache, a hive of its own And the Prefetch file itself
Persistence
Run and RunOnce keys, services, scheduled task registrations, Winlogon hooks, AppInit_DLLs, COM hijacks. Nearly every persistence technique has to write itself somewhere the OS will read at startup, which means the registry.
Devices
Every USB device ever attached leaves vendor, product, serial and first/last connection times under Enum\USB and Enum\USBSTOR, tied to drive letters through MountedDevices.
User activity
RecentDocs, OpenSave and LastVisited dialog history, typed paths, Explorer search terms, and Shellbags, which record folders a user browsed, including folders that no longer exist.
Network
Every network the machine joined, with first and last connection times, gateway MAC addresses and whether it was wired or wireless: a movement history for a laptop.
Time, on every key
Because each key carries a last-written timestamp, the registry is not just a state snapshot . It is a partially ordered record of when that state changed.
The catch this guide is really about
All of that assumes the hive file you are reading is the registry that was actually running. Very often it is not.
Windows does not finish a registry write in one go. It writes what it is about to do into a second file first, changes the hive, and only then marks the hive as settled. Copy a hive off a machine that is switched on and you will usually catch it mid-sentence: the file is complete and readable, but the most recent changes are not in it yet. They are in a file sitting right next to it, called SYSTEM.LOG1 or NTUSER.DAT.LOG1, which most tools do not open at all.
The practical result is that the newest registry activity, which is usually the activity closest to whatever you are investigating, is the part most likely to be missing. That is sections 15 to 17.
4. The Registry Is a Set of Files
Here is the step that turns the registry from an abstraction into evidence: the tree you see in regedit does not exist on disk. It is assembled at boot from a handful of separate files, each mounted at a name.
| Hive file | Location | Mounted at | Holds |
|---|---|---|---|
| SYSTEM | \Windows\System32\config\ | HKLM\SYSTEM | drivers, services, devices, mounted volumes |
| SOFTWARE | \Windows\System32\config\ | HKLM\SOFTWARE | installed applications, OS configuration |
| SAM | \Windows\System32\config\ | HKLM\SAM | local user accounts and group membership |
| SECURITY | \Windows\System32\config\ | HKLM\SECURITY | LSA policy, cached secrets |
| DEFAULT | \Windows\System32\config\ | HKU\.DEFAULT | the profile in force before anyone logs on |
| NTUSER.DAT | \Users\<user>\ | HKU\<SID> | one per user: settings and activity |
| UsrClass.dat | \Users\<user>\AppData\Local\Microsoft\Windows\ | HKU\<SID>_Classes | shell state, file associations, Shellbags |
Two consequences that matter constantly in casework.
HKEY_CURRENT_USERis not a file. It is whicheverNTUSER.DATbelongs to the account you are running as, mounted underHKEY_USERS. A user who is logged off has no entry there: their hive is on disk, unmounted, and invisible to any live tool.HKLM\SAMandHKLM\SECURITYcannot be read in place even as Administrator; their own ACLs deny it. They must be extracted with backup privileges or read from an image.
So: the tree is a view; the files are the evidence. Everything from here on is about what is inside one of those files.
5. Which File Backs Which HKEY
The previous section said the registry is a set of files. This one is the mapping, because the tree a tool shows you is not laid out like the files underneath it, and three of the branches you can browse are not in any file at all.
| What you see | What it is on disk | Notes |
|---|---|---|
| HKLM\SYSTEM | %SystemRoot%\System32\config\SYSTEM | services, drivers, device history, mounted devices |
| HKLM\SOFTWARE | %SystemRoot%\System32\config\SOFTWARE | installed software, most machine-wide persistence |
| HKLM\SAM | %SystemRoot%\System32\config\SAM | local accounts; the root key denies Administrators |
| HKLM\SECURITY | %SystemRoot%\System32\config\SECURITY | LSA secrets, policy |
| HKU\.DEFAULT | %SystemRoot%\System32\config\DEFAULT | the profile the logon screen runs under, not "the default user" |
| HKU\<SID> | C:\Users\<name>\NTUSER.DAT | one per user, mounted only while that user is logged on |
| HKU\<SID>_Classes | ...\AppData\Local\Microsoft\Windows\UsrClass.dat | per-user file associations, Shellbags |
| HKCU | no file | a link to HKU\<SID> of the calling user |
| HKCR | no file | a merged view of two Classes subtrees |
| HKLM\HARDWARE | no file | rebuilt in memory at every boot |
Three things that exist only while Windows is running
HKLM\HARDWARE is volatile. It is built by the kernel at boot from what it finds on the bus, and written to no file. If you are working from an image, it is not missing, it never existed on disk. The same is true of HKLM\SYSTEM\CurrentControlSet, which is a link created at boot to whichever ControlSet00n was selected.
HKCR is a merge, not a hive. It presents HKLM\SOFTWARE\Classes overlaid with the user's UsrClass.dat, with the user's entries winning. A COM hijack that appears in HKCR live may live in either file, and only one of them is per-user. When you work offline you must decide which you are looking at.
WOW6432Node is a redirection, not a location. On a live 64-bit system a 32-bit process asking for HKLM\SOFTWARE\Classes\CLSID is quietly served ...\Classes\Wow6432Node\CLSID. In the file, only the second path exists. A query written against the live illusion returns nothing offline, with no error to say why.
6. The Same Thing, Called Two Different Names
From here on this guide talks about the registry in two vocabularies at once, and it is worth stopping to say so, because almost every confusion about hive files comes from not noticing the switch.
One vocabulary is the one Regedit shows you: keys, subkeys, values, data. The other is what is actually in the file: cells, records, lists, offsets. They describe the same thing. Neither is more correct. But only the second one exists on disk, and a forensic tool reading a hive file has nothing but the second one to work with.
The translation table
Everything in the rest of this guide is one of these six things. When a later section says "the nk record", it means the thing Regedit would draw as a folder.
| What Regedit calls it | What it is in the file | What that actually means |
|---|---|---|
| The tree under one root | a hive | One file on disk, plus its .LOG files. HKLM\SYSTEM is a file called SYSTEM. |
| A key | an nk cell | One record holding the key's name, its timestamp, and the positions of everything attached to it. It does not contain its subkeys or values. |
| A key's subkeys | an lh, lf, li or ri list | A separate record, elsewhere in the file, holding the positions of the child keys. The parent only knows where that list is. |
| A value | a vk cell | Its name, its type, how long its data is, and where the data lives. |
| A value's data | a data cell, or inline | Just bytes. Nothing in them says what they mean; the type in the vk record is the only thing that does. |
| A key's permissions | an sk cell | Shared. Hundreds of keys with the same permissions all point at one copy. |
…
…
The same key, both ways
The table above is the translation. This is the same translation applied to one real key
out of the reference system's SYSTEM hive: on the left as Regedit would draw it, on
the right as it actually sits in the file. Select any row in either panel and its
counterpart lights up. Nothing here needs the byte-level sections that follow; it is here
so that when they say nk or lh, you have already seen what they
are pointing at.
Nothing connects directly
The three relations you actually care about are all indirect, and the middle hop is always a record Regedit does not show:
- A key to its subkey is two hops. The key's subkey list offset names a list, and an entry in that list names the child. The key never mentions the child directly.
- A key to its value is two hops, in exactly the same shape: the key names a value list, and an entry in that list names the
vk. - A value to its data is one hop, and sometimes none at all, because a small value keeps its data in the offset field itself.
That is why a key record is only 88 bytes while the key it represents can have thousands of children, and it is why a parser spends most of its time following offsets rather than reading records.
The dimmed rows on the right have no counterpart on the left: they are the subkey lists and the value list. Regedit never shows them and never mentions them, and they are most of the work a parser does. A key does not contain its children; it contains the position of a separate record that lists where the children are.
Two words this guide uses constantly
Both come up on nearly every page from here on, and both are worth pinning down now rather than guessing at.
An offset
A position in the file, counted in bytes from the beginning. Offset 4096 means "4,096 bytes in from the start of the file".
The registry format has no names inside it in the way a filesystem does. When one record needs to refer to another, it stores a number saying where that other record sits. That number is an offset, and following it means jumping to that position in the file and reading what is there.
Nearly every structure in this guide is mostly offsets. A key record is largely a small set of numbers saying where its name, its subkey list, its values and its permissions can be found.
A parser
Any program that reads the hive file directly, byte by byte, instead of asking Windows to read the registry for it.
Regedit is not a parser in this sense. It asks Windows, and Windows hands back a tidy tree. Every forensic tool is a parser, because the machine it is examining is usually switched off, or is not the machine the analyst is sitting at. Crow-Eye is a parser. So is every tool this guide mentions.
This matters because a parser has to make every decision Windows would have made for it, and the sections that follow are largely about the decisions that are easy to get wrong without any error appearing.
Why the difference is the whole point
Regedit shows you a tidy tree because Windows builds one for it. On disk there is no tree: there are records scattered through the file, pointing at each other by position. The tree is something a program assembles by following those pointers.
That is why deleted keys can be recovered, why a half-written hive can look perfectly healthy, and why two tools can read the same file and disagree. All three are covered later, and all three come from the same fact: what Regedit shows you is a reconstruction, not a picture of the file.
7. The Base Block
A hex editor shows a file as what it really is: a long row of numbered bytes, with no structure except the one you know to look for. Open a hive in one and the very first four bytes read regf. That is the format's signature, and checking it is the first thing any program reading a hive does, because if those four bytes are anything else the file is not a hive and nothing that follows will mean what you think it means.
The next 4096 bytes are the base block, which is this format's word for the file header. It is a fixed-size block of fields at fixed positions, and it answers four questions: which version of the format this is, whereabouts in the file the root key sits, how much of the file is actually in use, and, the one that matters most later, whether the file was closed cleanly or was still being written when it was captured.
Click any byte below to see what it is. This is a real SYSTEM hive from a running Windows 11 machine.
The checksum, and why it bites
The last segment of the map is the checksum at 0x1FC: d4 41 93 6d, or 0x6D9341D4 in this hive. It is an XOR of the 127 32-bit words before it. That phrase is doing a lot of work, so unpacking it: everything from the start of the file up to the checksum is 508 bytes, and reading those 508 bytes four at a time gives 127 numbers, each 32 bits wide. XOR them all together and you should get the value stored here.
Why XOR and not a real checksum. XOR is about as cheap as an operation gets, and this value is recomputed every time the header is flushed to disk, which happens constantly. It is there to catch a storage fault: a bit that flipped on the way to the platter, a sector that came back wrong. It is not there to catch anyone. Two bits flipped in the same position of two different words cancel out exactly, so anybody editing a hive deliberately can keep the checksum valid without effort. Treat it as a signal that the header is undamaged, never as a signal that it is unaltered.
Two results are reserved: if the XOR comes out 0 it is stored as 1, and if it comes out 0xFFFFFFFF it is stored as 0xFFFFFFFE. Any tool that writes a hive back, including the log recovery in section 18, must recompute it, or every reader rejects the file. Note also what the map makes obvious: everything from 0xB0 to 0x1FB is reserved and empty, so the header is 4096 bytes of which fewer than two hundred carry anything.
Two timestamps, and only one you can lean on
The last-written FILETIME at 0x0C reads zero in this hive. Across all seven hives on the reference system it is zero on six, and populated on UsrClass.dat, which records 2026-02-05 17:27:47. So it is written sometimes and not others. Check it; never assume it.
The one populated everywhere is last-reorganised, at 0xA8: when Windows last compacted the hive, rewriting free space and reordering cells. On the reference system all seven land within the same fifteen minutes of 2026-08-12: what a single scheduled maintenance pass looks like from the inside.
Neither is the timestamp forensics leans on. That is the per-key one in section 9, maintained on every key, and the closest thing the registry has to a clock.
What that path field really holds
0x30 is easy to read as a label and move on. Put every hive on one
machine side by side and it is plainly something else: the file's own path,
and on a user hive that includes the account name:
| Hive | 0x30 stored path | len | 0x0C written | 0xA8 reorganised |
|---|
Look at the lengths. SAM fits in exactly 31 characters;
SOFTWARE, SECURITY and DEFAULT are longer,
and every one of them has lost its beginning . The field keeps the
last 31 characters and drops the rest. That is why
SOFTWARE reads emRoot\System32\Config\SOFTWARE, missing
its leading backslash and the first four letters of "System".
Why this is worth knowing
A hive carries a record of where it came from. An NTUSER.DAT pulled
out of unallocated space, or renamed, or handed to you in a folder called
evidence\hive3.dat, still says
\??\C:\Users\Admin\ntuser.dat in its own header: naming the
account it belonged to, without any external context.
Treat it as a strong lead rather than proof: it is written when the hive is created, so a profile renamed afterwards keeps the old path, and nothing prevents the field being edited. But it is checked far less often than it deserves.
8. Hive Bins and Cells
After the base block, the rest of the file is one big memory allocator. Windows maps the hive into memory and manages space inside it the same way a heap manages RAM, which is exactly why the structures look the way they do.
The file after offset 0x1000 is divided into hive bins (hbin), each a multiple of 4096 bytes. Every bin begins with its own header: the signature hbin, its offset relative to the start of the bins area, and its size.
Cells sit end to end with no gaps and no index. You find the next one by reading the current one's size and stepping that far, which is why a single wrong size field makes the rest of the bin unreadable.
The same four bytes carry two completely different numbers depending on how you read them, and both readings are legitimate. Read a8 ff ff ff as a signed number and it is −88. Read the identical bytes as an unsigned number and it is 4,294,967,208.
Both are true, because a signed 32-bit number uses its top bit to mean "negative" while an unsigned one treats that bit as just another digit, worth two billion. Nothing in the bytes says which convention applies. The format says it, and the program has to know.
This is why getting it wrong is dangerous rather than merely wrong. A program that reads sizes as unsigned does not crash and does not complain. It reads this cell as four gigabytes long, adds four gigabytes to its position to find the next cell, lands far past the end of the file, and stops. It will report having read the hive successfully and will have read one cell of it.
Inside a bin sit cells: the unit of allocation, laid end to end. A cell's first four bytes are a signed 32-bit size, and the sign is doing real work:
Negative size = allocated. Positive size = free.
The root cell of this hive sits at file offset 0x1020 and its size field reads -88: 88 bytes, in use. Those four bytes are the reason the file can be read at all. There is no index of cells anywhere in a hive, and no table of contents. Every cell simply begins by stating its own length, so a program finds the next cell by reading the current one's size and stepping exactly that far forward. That is the entire navigation mechanism, and it is why one wrong size field makes everything after it in that bin unreadable: the program is no longer landing on cell boundaries.
On top of carrying the length, the sign carries the status. Negative means the cell is in use, positive means it is free. To work out how much space a cell actually gives you, take the absolute value and subtract the four bytes the size field itself occupies, so this 88-byte cell holds 84 bytes of record.
Freeing a cell therefore costs almost nothing: Windows flips the sign and moves on. It does not erase the contents, it does not shuffle anything up to close the gap, and it does not touch the bytes at all.
Two things follow from this design, and both matter forensically:
- Deleted data survives. Deleting a registry key marks its cell free by flipping the sign. The bytes, meaning the key name, its values and its timestamp, remain until something else allocates over them. This is why registry carving for deleted keys works at all.
- Everything is an offset. Cells refer to each other by their position relative to
0x1000, not by name. A hive is a graph of offsets, so a structure pointing at a cell that was never written is not a slightly damaged hive . It is an unreadable one. Hold on to that for section 15.
9. Key Nodes: the nk Record
A registry key is a cell whose contents start with the two letters nk. Everything the tree shows you about a key is in this one record: its name, its timestamp, how many subkeys and values it has, and where to find them.
| Offset | Size | Field | What it means |
|---|---|---|---|
| 0x00 | 2 | Signature | nk: this cell is a key node |
| 0x02 | 2 | Flags | root key, symlink, name is ASCII, and so on. This hive's root reads 0x2C |
| 0x04 | 8 | Last written | FILETIME, and this one is real. The timestamp forensics actually uses |
| 0x10 | 4 | Parent offset | the cell of the key above this one |
| 0x14 | 4 | Subkey count | stable subkeys: 17 for this hive's root |
| 0x18 | 4 | Volatile subkey count | subkeys that exist only in memory and are never written to the file |
| 0x1C | 4 | Subkey list offset | the cell holding the list of children |
| 0x24 | 4 | Value count | how many values this key holds |
| 0x28 | 4 | Value list offset | the cell holding an array of vk offsets |
| 0x2C | 4 | Security offset | the sk cell with this key's security descriptor |
| 0x48 | 2 | Name length | bytes of name that follow |
| 0x4C | n | Key name | not null-terminated: the length field is the only delimiter |
A trap worth naming
The subkey list offset is at 0x1C, not 0x18. 0x18 is the volatile subkey count. Reading the wrong one gets you a small integer where an offset should be, which then points somewhere meaningless inside the first bin, and a parser that does not validate will happily read garbage and report it as registry data. This exact mistake was made while writing this guide, and it was caught only because the resulting "subkey list" had a cell size of 7 235 938 bytes and a signature of \x00\x00. Validate signatures; do not trust offsets.
Volatile keys deserve a note of their own. Some registry keys exist only in memory and are never written to any file. HKLM\SYSTEM\CurrentControlSet itself is one, as is Control\hivelist. This is why a live registry and an offline hive legitimately differ: certain keys simply cannot exist in a file, no matter how the image was acquired.
10. Value Records: the vk Record
If nk is a folder, vk is a file. It holds a value's name, its type, how big its data is, and where the data lives, with one space optimisation that catches out anyone writing a parser.
| Offset | Size | Field | What it means |
|---|---|---|---|
| 0x00 | 2 | Signature | vk |
| 0x02 | 2 | Name length | 0 means the key's unnamed (Default) value |
| 0x04 | 4 | Data size | bottom 31 bits are the length; the top bit changes everything |
| 0x08 | 4 | Data offset | the cell holding the data: or the data itself |
| 0x0C | 4 | Data type | 1 = REG_SZ, 3 = REG_BINARY, 4 = REG_DWORD, 7 = REG_MULTI_SZ, … |
| 0x10 | 2 | Flags | bit 0 set = the name is ASCII rather than UTF-16 |
| 0x14 | n | Value name | again, length-delimited, not null-terminated |
The inline-data trick
Normally a value's vk record does not contain its data. It contains a number saying where the data is, and the program reads that number and jumps to that position in the file to fetch it. That jump is what "following an offset" means, and it happens millions of times in a registry parse.
For very small values that is wasteful. Allocating a whole separate cell to hold four bytes costs more in bookkeeping than the four bytes are worth. So the format has a shortcut: if the top bit of the data-size field is set, the data is not somewhere else at all. It is sitting in the four bytes that would otherwise have held the position. The field that normally says "the data is over there" instead says "the data is right here, and it is this".
The trap is that nothing about the record looks different. A program that always treats that field as a position will, for a REG_DWORD whose value happens to be 1, jump to position 1 in the file, read whatever bytes are sitting in the header there, and report them as the value's data. The length will be right. The type will be right. The data will be nonsense, and no error will be raised anywhere, because from the program's point of view it did exactly what it was told.
The fix is one line and one branch: strip the top bit off before using the length, and use it to decide whether to read the field as data or as a position.
TypeID at 0x16A0, cell wrapper included.The explorer in section 8 walks to a real example: the value TypeID under ActivationBroker\Plugins, 78 bytes of REG_SZ data, stored in its own cell because it is far too big to inline.
11. Value Types, and Data Too Big to Fit
The vk record carries a type and a length. Both have a trap in them. The documented set means the twelve REG_* type constants that Microsoft publishes and the format specification records, and this field is not limited to them. The length field, meanwhile, has a ceiling above which the data stops being in one piece.
The documented types, and how often they are actually used
Counted across the three hives on the reference system, every allocated vk record, … values in total:
| ID | Name | What the bytes are | Count |
|---|
… REG_SZ and REG_DWORD between them account for most of everything.
Types outside the twelve documented ones
… values on the reference system carry a type that is not one of the twelve. They are not corruption, and they are not spread at random: every one of them sits under a device or driver key, in Enum, DeviceClasses, DriverDatabase\DriverPackages or Setup\Upgrade\PnP.
The data itself says what they are. A value typed 0xFFFF0012 holds UTF-16 text; 0xFFFF0007 is always exactly four bytes; 0xFFFF0011 is one byte; 0xFFFF000D is sixteen, the size of a GUID; and 0xFFFF0010 is eight bytes that decode as a FILETIME. That is the device property type space, which Plug and Play stores in the same field.
What it means for a parser: a switch on the type field with twelve cases and no default will silently drop thousands of values on a normal Windows install, and every one of them is device history.
When the data does not fit
A value's data normally lives in one cell. Past a certain size it cannot, and Windows stores a db record instead: a small cell holding a count and a pointer to a list of segment offsets, with the real data split across those segments in order.
The boundary is easier to measure than to look up. Across these three hives the largest value still held in a single cell is … bytes, and the smallest one that needed a db record is …. The real limit sits between the two, which is what you would expect: a cell cannot cross a hive bin boundary, and a bin is 4096 bytes times some whole number.
…
The reason for the split is the allocator from section 8. A cell lives inside one hive bin and may not straddle the boundary into the next, so no single cell can be larger than a bin, and a value bigger than that has to be broken up. The db record is how the format keeps the vk record's shape unchanged while the data behind it becomes a list.
What this means for a parser. The vk record looks exactly the same either way. Same signature, same length field, same offset field. The only thing that distinguishes a value with 200 KB of data from one with 20 bytes is what is sitting at the far end of the offset, so the check has to be made every time: read the two bytes at the target and see whether they are db. Skip that check and every large value on the system returns the same wrong thing, silently, and it will be exactly the right length.
These hives hold … of them. A parser that reads the data offset and takes the next length bytes returns the first segment followed by whatever happened to be allocated next, which is a plausible-looking string of exactly the right length and the wrong content.
12. Subkey Lists
An nk record does not contain its children. It contains the offset of a list of them, and there are four kinds of list, because the registry has to stay fast when a key has tens of thousands of subkeys.
| Signature | Name | Entry | Purpose |
|---|---|---|---|
| lf | Fast leaf | offset + first 4 name characters | early format; compare the hint before following the offset |
| lh | Hash leaf | offset + name hash | the modern form: this hive's root uses it |
| li | Index leaf | offset only | no hint; used where hints do not help |
| ri | Index root | offsets of other lists | a list of lists, for keys with more children than one leaf holds |
The hint stored beside each offset is the point. To find a subkey by name, Windows compares the cheap hint first and only follows the offset when the hint matches, because following it means a pointer chase into a different part of the mapped file. Entries are kept sorted, so lookup is a binary search rather than a scan.
ri: a list of lists…
For anyone writing a parser: lf and lh have 8-byte entries, li has 4-byte entries, and ri entries point at further lists that must be walked recursively. Assuming an 8-byte stride for all of them reads every other entry as garbage on an li list, and produces a subkey count that looks plausible.
This is the row the map in section 6 dims out. A key does not contain its children: it contains the position of one of these lists, and the list contains the positions of the children. Regedit shows you neither.
All of it, inside each other
Each map above shows one structure flat. What they cannot show is that every one of them lives inside another, and that the way you get from one to the next is by following an offset out of the record you are looking at. That is what a parser does, and it is what this walks.
Two things worth noticing as you walk it.
- The cell wrapper is always there. Every record begins with four bytes that are not part of the record. They belong to the allocator. "Cell" and "nk" are not two words for the same thing; one contains the other.
- Nothing is nested in the file. A key's children are not stored
inside it; they are somewhere else entirely and it holds their offset. The
hierarchy you see in
regeditis reconstructed by chasing those offsets, which is exactly why a single broken one makes a hive unreadable.
13. Security Descriptors and Class Names
Every pointer out of an nk record has now been followed except one. The security offset leads to an sk record, and it is the reason a registry with hundreds of thousands of keys is not enormous.
One descriptor, shared by hundreds of keys
An sk cell holds a Windows security descriptor: owner, group, and the ACL that decides who can read or write the key. What makes it interesting is that keys do not get one each.
Consider the arithmetic if they did. This SYSTEM hive has 69,624 keys, and a modest security descriptor runs to a few hundred bytes. One descriptor per key would be tens of megabytes of almost entirely duplicated data in a file that is 23 MB in total. So the format does the obvious thing: identical descriptors are stored once and shared, each carrying a count of how many keys are using it, and the records are threaded together so Windows can walk them when it needs to find an existing match rather than allocate a new one.
| Hive | Keys | sk records | Keys per descriptor | Keys with a class name |
|---|
sk ring…
Why an examiner should care
A key whose permissions differ from its siblings has its own sk record, and that is visible structurally without reading a single ACL. Where thousands of keys share one descriptor, an outlier is worth a look: weakened permissions on a persistence key is a way to let a non-administrator write to it, and it leaves no trace anywhere else.
Note also what is not here. The sk record stores SIDs, not names. Resolving them against the machine you are sitting at is how an offline account gets confidently mislabelled as a local one.
Class names, the field almost nothing shows you
An nk record can carry a class name, a second string separate from the key's name and stored in its own cell. Most keys have none. On the reference system … do, and in one place the field is not a label at all: it is where the data lives.
The whole of SYSTEM has just … keys with a class name, and they are these:
| Key | Class name length |
|---|
Four of those five are the boot key
JD, Skew1, GBG and Data under ControlSet001\Control\Lsa each hold eight hexadecimal characters in their class name, and nothing in their values. Concatenated in a fixed order and then permuted, those thirty-two characters are the SysKey, the machine's boot key, which is the input to decrypting the password hashes in the SAM hive.
This is the reason the field is worth knowing about. Windows put a secret somewhere that most registry viewers do not render, and a tool that reads names and values and ignores class names cannot see it at all, while the bytes sit in plain text in the file.
The values are not printed on this page. Every other figure here comes from a real machine and is shown; these four are the exception, because together they are that machine's boot key. The lengths are shown instead, and they are the part that makes the point.
Elsewhere the field is mundane: 535 keys in this NTUSER.DAT carry the class Shell, and a thousand Classes\CLSID keys in SOFTWARE carry REG_SZ. Mundane or not, it is a place data can be and most tools will not look.
14. Carving Deleted Keys and Values
The allocator section made the claim that deleted registry data survives. This section is the number behind it, because a claim like that is worth very little without one.
Deleting a key does not erase it. Windows flips its cell's size field from negative to positive, marks it free, and moves on. The signature, the name, the timestamp and the pointers are all still there, and stay there until something else allocates over that space.
That single fact opens a second way to read a hive. Every tool you have used walks the tree: start at the root, follow subkey lists down, and report what you reach. A key that has been unlinked from the tree is invisible to all of them, and it is still in the file. The other way is to ignore the tree entirely and walk the allocator: step through the bins, read each cell's size field, and look at everything the tree can no longer reach. That walk is what the panel below does, one cell at a time.
| Hive | Bins | Allocated cells | Free cells | Free space | Still an nk | Still a vk |
|---|
…
What comes back out
These are real keys, read out of free space in the reference system's SYSTEM hive. Nothing in the tree points at any of them:
| Cell | Recovered key name | Last written | Values |
|---|
A deleted key keeps its own last-written timestamp, which is often the more useful half: it dates the activity, not the deletion.
What a carved key is and is not
It is not proof of deletion by a person. Windows frees registry cells constantly by itself: uninstallers, driver updates, profile maintenance and the periodic reorganisation from section 7 all leave freed cells behind. A key in free space means it was removed, not that anyone removed it deliberately.
Its parent is a pointer, not a path. A recovered nk holds its parent's offset, and that cell may itself have been reused. Where the chain breaks, the key's full path is unknown, and inventing one is how a carved artefact turns into a wrong conclusion. Following those pointers is still worth doing, because it works more often than not: in the reference system's SOFTWARE hive, 941 of the 1,451 carved keys walk their parents all the way back to a key that is still live and get a full path out of it. The other 510 keep their name, their timestamp and their values and are recorded with no path at all, because a partial path reads exactly like a real one.
Free does not mean intact. The counts above are cells whose signature still reads nk or vk. Part of a record can be overwritten while its first two bytes survive, so every field has to be sanity-checked before it is believed: a name length of 60,000 or a timestamp in the year 30,000 means you are reading someone else's data.
Nothing in the record says it was deleted. A carved nk is byte for byte the same shape as a live one; the entire claim rests on where the reader found it. An allocator walk that miscounts a single cell size, or a pointer followed back into space that is still in use, produces a table of deleted keys that were never deleted, and no field anywhere in that output would look wrong. The claim has to be checked from outside itself: walk the tree separately, and require every carved offset to be absent from it.
The reorganisation timestamp matters here too. Compacting a hive rewrites its free space, so a hive that was reorganised recently has less recoverable history than one that was not, and the field at 0xA8 tells you which you are holding before you spend an afternoon on it.
15. Why a File Like This Needs a Journal
That is enough structure to see the problem. Everything above is offsets pointing at other offsets, and a single logical change touches several of them at once.
This is the hinge of the guide. Everything before it describes what a hive looks like when it is intact. Everything after it exists because a hive on disk is often not intact, and the parts that say so are the parts most tools skip.
Take something trivial: an application writes a registry value. Follow it through the structures from sections 8 to 12:
- The new data does not fit the existing cell, so a new cell is allocated somewhere else in a bin.
- The
vkrecord's data offset is updated to point at it. - The old cell's size is flipped positive to free it.
- The parent
nk's last-written timestamp changes. - If a bin had no room, a new bin is appended and the base block's bins-size updated.
Those are separate writes to separate pages of the file, and they only make sense together. Interrupt them, whether by power loss, a forced reset or a snapshot taken at the wrong instant, and some pages reached the disk while others did not.
The important thing is what that failure looks like from the outside, which is: fine. A hive is not a text file, where a truncated write leaves an obviously ragged ending. Every structure in it is self-describing and independently valid. A vk record whose data offset points at a cell that was never written still has a correct signature, a plausible name length and a type in range. A parser reads it, follows the offset, and returns whatever bytes happen to be sitting at that address. It reports a value. It does not report an error, because from where it is standing nothing is wrong.
Nor does the base block help. Its checksum covers the first 508 bytes of the header and nothing else, so a hive whose body is half-updated still passes the only integrity check the format has. That is the situation the rest of this guide is about: a file that parses cleanly and describes a state the machine was never in.
Two consequences follow, and they are the reason the second half of this guide exists. The first is that Windows needs a way to finish an interrupted write after the fact, which means writing its intentions down somewhere else before it starts. The second is that an examiner needs a way to tell whether the file in front of them is the finished state or the half-finished one, and that needs a flag.
So the registry does what databases and journalling filesystems do: it writes down what it is about to do, somewhere else, first. That somewhere else is the transaction log, and the flag that says whether it is needed is a pair of counters in the base block.
16. The Two Sequence Numbers
You saw them in the base block map at 0x04 and 0x08, coloured red. Here is why there are two of them and not one, and it is worth playing with the control below rather than just reading it.
The write protocol has three steps:
One comparison, and you know.
Sequence1 == Sequence2 → the hive was closed cleanly.
Sequence1 != Sequence2 → a write was in flight; the log holds the rest.
This is what every tool means when it says "dirty hive". It is not a heuristic . It is two integers and an equality test. A single counter could say how many writes had happened, but never whether the current one had finished.
What it looks like on a machine that is switched on
Seven hives, captured from one Volume Shadow Copy snapshot of a running Windows 11 system:
| Hive | Sequence1 | Sequence2 | State |
|---|---|---|---|
| SYSTEM | 1,888,230 | 1,888,229 | DIRTY |
| SOFTWARE | 5,102,671 | 5,102,670 | DIRTY |
| NTUSER.DAT | 365,687 | 365,686 | DIRTY |
| UsrClass.dat | 251,613 | 251,612 | DIRTY |
| DEFAULT | 197,479 | 197,478 | DIRTY |
| SAM | 382 | 382 | clean |
| SECURITY | 1,069 | 1,069 | clean |
Five of seven are dirty, and every dirty one is dirty by exactly one. That is not five interrupted writes . It is the ordinary steady state of a machine that is running, because there is always another write in flight. The two clean ones are the two nobody writes to: SAM has been written 382 times in the life of this installation, while SYSTEM has been written 1.8 million times.
For an examiner
On a live-acquired image, expect the interesting hives to be dirty. It is the normal condition, not damage, and not evidence of anti-forensics.
17. The Transaction Logs
Next to every hive sit two more files with the same name and the extensions .LOG1 and .LOG2. This is where the outstanding changes live, and where almost every forensic tool stops.
Two formats
The old format, up to Windows 7, marks its dirty data with a DIRT signature. The new format, from Windows 8.1 onwards, is the one described here. A parser must be able to recognise the old one in order to refuse it. Reading a DIRT log with new-format offsets does not raise an error, it reads plausible bytes and writes them into the wrong parts of the hive.
A new-format log file looks like this:
| Offset | Contents |
|---|---|
| 0 | a 512-byte copy of the hive's base block: with its own sequence number |
| 512 | the first log entry (HvLE), then the next, and the next |
| … | unused space, which is not empty. See the stale tail below. |
Note the base block copy at the front, and that it carries its own sequence number. That single detail is the crux of the investigation at the end of this section, and the difference between a working recovery and a corrupted hive.
Inside a log entry
After the 40-byte header comes a dirty page reference table: one 8-byte pair per page, an offset and a size, and then the page data itself, back to back in the same order. To apply an entry you walk the table and, for each pair, write the next size bytes into the hive at 4096 + offset.
These are page images, not deltas
Every offset is 4096-aligned and every size is a multiple of 4096. An entry does not say "change these bytes". It says "this page now looks exactly like this". Measured across the 53 entries of one real log: 620 page writes to only 149 distinct offsets, an average of 4.2 rewrites of the same page, with offset 0, the first bin, holding allocation bookkeeping, present in all 53.
So a later entry completely supersedes an earlier one for the same offset. That is why entries must be applied in sequence order, last write wins.
Why a signature is not enough
A log file is not truncated when it is reused. Windows writes entries from offset 512 onward and leaves whatever was there before in place beyond the last one. The tail of a log file is the previous generation's entries: real, structurally valid, correctly hashed entries that are simply old.
A parser that scanned for HvLE and applied everything it found would apply 9 stale transactions to SYSTEM and 15 to UsrClass.dat, writing an older generation's pages over a newer hive. It would not crash. It would produce a hive that opens cleanly and is quietly wrong.
Two mechanisms prevent it. The sequence chain: entries run consecutively, and parsing stops at the first entry whose sequence is not the one expected next, so the stale tail is never reached. Marvin32: each entry carries two hashes, one over its body from offset 40 and one over its first 32 bytes, seeded with 0x82EF4D887A4E55C5. Both must reproduce or the entry is not an entry.
If you are implementing this: reproducing those two hashes is the best self-check available. If your Marvin32 and your header offsets are both right, the stored values come out exactly; if either is wrong, they do not. It validates the algorithm and the layout in one step, before you have written a single byte to a hive.
Why there are two logs
Windows alternates between them, so that at any moment one holds a complete, applicable run. Together the pair covers one unbroken sequence range:
| File | Entries | Range | Pages |
|---|---|---|---|
| SYSTEM.LOG1 | 53 | 1889322 → 1889374 | 620 |
| SYSTEM.LOG2 | 34 | 1889375 → 1889408 | 535 |
Which file comes first is not fixed
For UsrClass.dat on the same machine it was the other way round: LOG2 covered 251897–251915 and LOG1 continued 251916–251932. Order the pair by the sequence numbers inside them, never by filename. Assuming .LOG1 leads works on most hives and silently produces the wrong result on the rest.
18. A Gap That Should Not Be There
Everything above is the format as it is understood now. This section is how one part of it was actually established, because the documented rule and the real bytes disagreed, and the two possible explanations demanded opposite implementations.
It is worth reading slowly. This is the hardest thing in the guide, and it is here because the answer is less useful than the method: what to do when the documentation and the file cannot both be right.
The recovery rule as commonly described is that log entries continue from where the hive left off. For SYSTEM, with Sequence2 at 1,888,229, the first log entry should have been 1,888,230. It was 1,889,322. Measured from the hive's own position, that first entry sits 1,093 sequences ahead of it. And it was not a one-off:
| Hive | Sequence2 | First log entry | Distance first entry − Sequence2 |
|---|---|---|---|
| SYSTEM | 1,888,229 | 1,889,322 | 1,093 |
| SOFTWARE | 5,102,670 | 5,104,537 | 1,867 |
| NTUSER.DAT | 365,686 | 366,244 | 558 |
| UsrClass.dat | 251,612 | 251,897 | 285 |
| DEFAULT | 197,478 | 197,940 | 462 |
The first suspicion was acquisition: three files copied at three different moments would explain it. They were re-copied from a single snapshot, together. Identical numbers, and the base block checksums validated. This was the real on-disk state.
Reading A: the transactions are gone
Everything between the hive's position and the log's start has been overwritten. Applying these entries writes page images onto a state they were never computed against. A replay must refuse.
Reading B: the base block lags
The hive's body is further along than its header admits, and the distance is an artefact of measuring from the wrong place. A replay must apply.
Reading A: the transactions are gone
Everything between the hive's position and the log's start has been overwritten. The two states do not meet, and applying these entries writes page images onto a state they were never computed against, producing a file that looks valid and is subtly wrong.
→ A replay must refuse.
Reading B: the base block lags
The hive's body is further along than its header admits, the log continues correctly from the real position, and the distance is an artefact of measuring from the wrong place.
→ A replay must apply.
One of these corrupts evidence. There was no way to tell from the file alone which was true.
A hypothesis, tested and discarded
Reading B implies the hive's body is ahead of its base block, which is testable, because the base block declares the hive bins size at 0x28. If the body had grown past what the header records, the file would be bigger than the header claims. It is: SYSTEM holds 23,851,008 bytes of bins where the header says 23,826,432.
Encouraging, until the log entries were checked. The last log entry declares the same bins size as the base block, not the larger figure. If the log's own final view of the hive agrees with the header, the surplus is not unrecorded growth; it is space Windows pre-allocated and has not used. The hypothesis predicted disagreement and found agreement, so it was wrong.
Why guessing was not an option
It is worth being precise about the stakes, because they are asymmetric and neither error announces itself.
Implement Reading A and you refuse to replay logs that were perfectly good. The tool reports the hive as it was found, the analyst never learns that more recent registry activity existed, and nothing anywhere says a decision was made. Evidence is quietly left on the floor.
Implement Reading B and you replay logs that should have been refused. Page images computed against one state are written over a different one, and the output is a hive that parses cleanly, passes its checksum, and contains a mixture of two moments in time. Evidence is quietly manufactured. This is the worse of the two, because the result looks like a successful recovery.
So a coin flip was not available, and neither was picking the reading that was easier to implement.
How it was settled
At that point the specification, the file, and reasoning about it were exhausted. The remaining options were to guess or to find something that already knew, so the implementation was checked against yarp, the reference implementation of registry log recovery, written by the author of the format specification and sharing no code with it.
It recovered all five hives without complaint. That answers the question empirically, and reading how it decides reveals the rule the common description glosses over:
The rule
Log entries chain from the log file's own base-block sequence number, not from the hive's. A log is eligible when its sequence is at or ahead of the hive's Sequence2; only a log older than the hive is refused.
For SYSTEM.LOG1 that number is 1,889,322, exactly its first entry. The log is self-describing. It was never claiming to continue from the hive; it states where it begins.
The gap was never a gap. It was a chain being measured from the wrong end.
Agreement with a reference is not the same as being right, so the bar was set higher: for every hive, the two implementations had to produce a byte-identical recovered file, or refuse for the same reason. They did: five recovered identically, both declining the two clean hives, with every source file's SHA-256 unchanged either side of the run.
What it recovers
| Hive | Entries | Pages | Keys gained | Values gained | Lost |
|---|---|---|---|---|---|
| SYSTEM | 87 | 1,155 | 0 | +5 | 0 |
| SOFTWARE | 250 | 3,905 | +4 | +4 | 0 |
| NTUSER.DAT | 69 | 1,082 | +10 | +4 | 0 |
| UsrClass.dat | 36 | 188 | +1 | +2 | 0 |
| DEFAULT | 17 | 233 | 0 | 0 | 0 |
Ten keys and four values in one user's NTUSER.DAT: the most recent activity there is, and exactly the part of a user hive an examiner cares about most. Without replay it is not "slightly stale"; it is absent, with nothing to indicate anything is missing.
What is still not known
Why does Windows leave the base block hundreds or thousands of sequences behind its log? Reading B is confirmed by behaviour, but the mechanism is still open. The candidates are that the log is reinitialised at a point which does not update the primary's header, or that the primary is flushed on a schedule independent of the base-block sequence. Nothing measured here distinguishes them, and one hypothesis in that family has already been tested and failed.
Settling it would need instrumented hive flushes on a live system, or kernel-side tracing of the registry writer. Neither is available from a snapshot. The recovery is verified; the explanation for one of its inputs is not, and saying so is more useful than a confident guess.
19. What a Hive Supports, and What It Does Not
The structure tells you what changed, never who changed it
Nothing in a hive file records an identity. Not the base block, not a key node, not a value record, not the transaction logs. A sequence number tells you a write happened and in what order; a key's timestamp tells you when the key was last touched. Which account, which process and which session did the touching lives outside the registry entirely.
A recovered cell is data, not a guaranteed record
Carving works because a deleted cell is unlinked rather than erased. That also means the bytes you recover may be a complete value, a partially overwritten one, or the tail of something else that happens to parse. A carved key with a plausible name and an implausible timestamp is usually the second case. Recovery from free space earns a lower confidence than a live key, and it should be reported as such.
Two parsers disagreeing is information, not a bug
A hive with unmerged transaction logs has two legitimate readings: the file as it sits, and the file as Windows would see it after recovery. Tools differ on which one they show, and neither is wrong. When two tools disagree about a registry finding, the logs are usually the reason, and the difference between the readings is itself the evidence of a write that had not yet landed.
A timestamp covers a key, not a value
Key nodes carry a last-write time; value records carry none. Every value under a key shares that one timestamp, and it moves when any of them is written, when a subkey is created, and when a value is deleted. It is an upper bound on everything beneath it — never a time of use for one particular setting.
20. What This Changes in Practice
Collect the logs
.LOG1 and .LOG2 belong in an acquisition beside every hive. A collection that takes only the hive files has discarded the most recent registry activity: the part closest to whatever prompted the investigation.
Say whether it was recovered
Whether a hive was replayed changes what its rows mean. A tool should record, per hive, whether it was dirty, whether logs were present and whether recovery ran, and when it cannot recover, parse anyway and say so rather than fail silently in either direction.
Never write to the evidence
Recovery produces a modified hive by definition. It belongs in a working copy, with the original opened read-only and its hash unchanged before and after.
One snapshot, all the files
A hive and its logs must come from the same acquisition instant. Copy them at three different moments and you will spend a day chasing an inconsistency that you created, as nearly happened here.
The door you come through decides what is in the file
The hives are locked while Windows runs, which is why everything on this page was collected from a shadow copy. It is worth being exact about what the alternative cannot do, because none of it is a matter of privilege. The registry API has no concept of a freed cell, so nothing in section 14 is reachable through it at any privilege level that exists. And some keys deny even an elevated administrator outright: the roots of SAM and SECURITY, and every device Properties subkey. An API walk of those simply returns a smaller tree, with no error anywhere in it. On the reference system an elevated walk of Enum\USB reaches 110 keys and is refused 21 subkeys; read as a file, the same hive gives 868.
So the answer is to stop asking the API and take the file instead, which three routes do: an existing Volume Shadow Copy, raw access to the disk underneath the lock, and a backup-privileged export, where you ask Windows itself to write the key out for you.
Those three are not equivalent, and the difference lands squarely on the two sections above. A shadow copy or a raw read hands you the hive as it sits on disk, allocated cells and freed cells alike. An export is a fresh file: Windows walks the live tree and writes out what it finds, and a freshly written hive has no free space in it at all. Carve an exported hive and you recover nothing, which is a true statement about that file and a false one about the machine it came from. Which door a hive arrived through is therefore part of what it is, and belongs recorded beside it, because it decides what an empty result means.
Reproduce all of it
Every byte on this page came from files any Windows machine has. The hives are locked while Windows runs, so they must be read through a Volume Shadow Copy or raw disk access. From there: the sequence numbers are two uint32 at 0x04 and 0x08; the first hive bin is at 0x1000; the root cell is at 0x1000 + the offset in 0x24; the log's first entry is at offset 512. A hex editor is enough to confirm every table here.
21. Why One Registry Needs Several Parsers
A reasonable question, having got this far: if every artifact in this file is a key with values, why does a forensic tool not have one registry parser? Crow-Eye has four that touch the registry, and the reason is not organisational tidiness. Each one exists because a different part of the problem is genuinely different.
The first split is the door, not the data
The same key can be reached two ways, and they are not the same operation.
| Route | Door | Sees | Cannot see |
|---|---|---|---|
| Live | Windows API | the merged tree, volatile keys, CurrentControlSet, the redirected 32-bit view | anything about the file: sequence numbers, unwritten log entries, deleted cells |
| Offline | the hive file | every byte in this article, including the parts Windows would have hidden | volatile keys, which were never written to a file at all |
Everything in the second half of this page follows from that difference. A live read cannot tell you a hive was dirty, because there is no hive from its point of view. An offline read cannot show you HKLM\SYSTEM\CurrentControlSet, because that key is assembled at boot and never stored; the file has ControlSet001 and a Select\Current value saying which one is in force. Crow-Eye resolves that itself, and on the reference system Select\Current reads 1, so the offline parser reads ControlSet001.
The two doors disagree, quietly
The library behind an offline read reports a key's unnamed value as the literal string (default), lowercase. The Windows API reports it as an empty name. Nothing errors either way, so the same key produces a different row depending on which door it was read through, and any comparison between the two silently never matches. That is the shape of nearly every live-versus-offline disagreement: not a crash, a difference.
The second split is what is inside the value
This is the one that actually forces separate code, and it is the point section 11 was building to: a value's data is just bytes, and nothing in the registry says what they mean. The type field tells you it is REG_BINARY. It does not tell you the blob is a shell item ID list.
Nothing to decode
A REG_SZ holding a service's image path is finished the moment it is read. Most of the registry is like this, which is why one parser covers most of it.
A structure from somewhere else
A Shellbag value is a shell item ID list: the same structure Windows writes inside a .lnk file. Decoding it is not registry work at all, which is why that code is shared with the shortcut parser rather than living in the registry one. Explained in full: the shell item And the shortcut it also lives in
An entire database
ShimCache is one value holding a serialised store with its own header, version signature and record layout. On the reference system that single value is 282,902 bytes. The registry is only the envelope it arrived in. Explained in full: one ShimCache record
A type Windows never documented
Device keys store Plug and Play property types in the same field as the twelve REG_* types. A reader that only knows the twelve refuses them, and one refused value is enough to lose the key it sat in.
So the four parsers divide along the two axes above, not by artifact:
| Parser | Opens | Exists because | Cannot be folded into the others |
|---|---|---|---|
| Live registry | the API | the door is the Windows API | it can see volatile keys nothing else can |
| Offline hive | the file | the door is the file | it can see hive state, logs and deleted cells nothing else can |
| Binary value | a blob | the payload is a structure the registry knows nothing about | it is called by both doors, and by the shortcut parser too |
| ShimCache | one value | the payload is a database, not a value | its record layout changes with the Windows version, not with the hive format |
Which is the argument for reading a file format at this depth in the first place. A parser is a claim about a file. The only way to know whether the claim holds is to open the file yourself.
22. Sources and Credit
This page describes a file format that Microsoft has never published. It is legible at all because other people did the work of writing it down, and the honest thing is to say whose work.
Maxim Suhanov, the format specification
The Windows registry file format specification at github.com/msuhanov/regf is the documentation this guide is checked against. The base block layout, the cell and record structures, the four subkey list types and the transaction log formats are all described there in detail, and where this page states an offset with confidence rather than hedging, that is why.
It is also unusually honest documentation: it marks what is uncertain as uncertain, which is what made it possible to tell the difference between a rule that was wrong and a rule this page had misread.
yarp, the independent check
Yet Another Registry Parser, by the same author, at github.com/msuhanov/yarp. It was used as an oracle for section 18: an implementation of log recovery sharing no code with the one being tested, so agreement between them meant something.
It is a development-time tool here and nothing more. It is not in requirements.txt, it is not imported by anything that ships, and no part of the product depends on it. Its job was to answer one question and it did.
The hives themselves
Every number, byte map and table on this page was measured from real hives captured from a running Windows 11 machine through Crow-Eye's own volume shadow copy path, alongside their .LOG1 and .LOG2 files. None of it is an illustrative example, and none of it was typed in by hand: a generator reads the hives and emits the data the page renders, and a separate script re-measures every published figure and fails if the page and the files disagree.
That is also the limitation. These are seven hives from one machine. Where a statement is true of these files rather than of all Windows systems, the page tries to say so, and where something was tested and turned out not to hold, it says that too.