Forensic Masterclass

The Windows Registry, End to End

From "what is a registry key" to the transaction log that recovers a dirty hive

Scroll to investigate

Hey Investigator,

Almost every Windows forensic finding passes through the registry at some point: what ran, what persisted, which USB stick was plugged in, which network the machine joined. Yet most investigators meet it through a tool that shows a tree, and never see what is underneath.

This guide takes it apart completely. It starts with what the registry actually is and why Windows cannot function without it, then goes down to the bytes: hive files, the base block, cells, the records that are keys and values, and finally the transaction logs: the part almost every tool ignores, and the reason a hive on disk is often not the registry that was really running.

No prior knowledge assumed. Every byte shown below was read out of a real SYSTEM hive captured from a running machine, so you can follow along in a hex editor.

Where these figures come from. Every number, byte map and table on this page was measured from a live Windows 11 system, its hives and transaction logs captured together through a single Volume Shadow Copy. That system is called the reference system throughout. Identifying values are redacted; every offset, length, count and timestamp is the real measurement.

1. What Is the Registry?

Before any byte-level work, one shared picture: the registry is a hierarchical database of configuration. Programs, drivers and Windows itself store settings in it instead of in their own scattered files.

Think of it as the building's central records office. Every department could keep its own filing cabinet in its own room, and once did, but then nothing can be indexed, secured, or backed up as a unit.

What came before, and why it failed

Early Windows kept configuration in .INI text files: WIN.INI, SYSTEM.INI, and one per application. That approach broke down for reasons worth naming, because the registry is a direct answer to each:

No data types

An INI file stores text. Everything, whether a number, a flag or a binary blob, arrives as a string, and every program has to parse and validate it itself. The registry has typed values: REG_DWORD, REG_SZ, REG_BINARY, REG_MULTI_SZ and others, so the type travels with the data.

No security

An INI file's protection is the file's own permissions: all or nothing. The registry attaches a security descriptor per key, so one branch can be writable by an administrator and readable by everyone, and Windows can enforce it.

No atomicity

Interrupt a program writing an INI file and you get half a file. The registry is transactional: the subject of most of this guide, so an interrupted write can be detected and undone.

No timestamps

An INI file has one modification time for the whole file. The registry stores a last-written timestamp on every key, which is the only per-record time it keeps, and the backbone of registry timelining. Read it carefully though: it records that something about that key changed, not what. See section 9.

The shape of it

The registry is a tree. Keys are the folders; they can contain other keys. Values are the leaves; each has a name, a type and data. A path like HKLM\SYSTEM\CurrentControlSet\Services\Tcpip is a chain of keys, and inside the last one sit values such as Start or ImagePath.

The names you know, such as HKEY_LOCAL_MACHINE and HKEY_CURRENT_USER, are called root keys or hives-in-the-abstract. section 4 shows what they really are: mount points onto files.

2. Why Windows Cannot Function Without It

It is tempting to think of the registry as "settings": the things you change in a preferences dialog. It is far more load-bearing than that. Windows reads the registry before it has a desktop, before it has services, and before most of the operating system exists.

Consider what the kernel needs to know at boot, before anything else can happen:

The consequence: the SYSTEM hive is loaded by the boot loader itself and handed to the kernel in memory. It is not a file the operating system opens once it is running . It is one of the things that makes it run. A corrupt SYSTEM hive is not a misconfigured machine; it is a machine that does not boot.

This is also why the registry needs the crash-safety machinery in sections 15 to 17. A configuration file that can be half-written is an annoyance. A boot-path database that can be half-written is an unbootable computer.

3. Why It Matters in Digital Forensics

Logs record what a system chose to write down. The registry records what a system was configured to do, and it keeps that record whether or not anyone intended it as evidence.

Execution

UserAssist counts and times GUI program launches. BAM and DAM record executable paths with last-run times per user SID. Prefetch's configuration, ShimCache and AmCache all live in or beside the registry. Explained in full: AmCache, a hive of its own And the Prefetch file itself

Persistence

Run and RunOnce keys, services, scheduled task registrations, Winlogon hooks, AppInit_DLLs, COM hijacks. Nearly every persistence technique has to write itself somewhere the OS will read at startup, which means the registry.

Devices

Every USB device ever attached leaves vendor, product, serial and first/last connection times under Enum\USB and Enum\USBSTOR, tied to drive letters through MountedDevices.

User activity

RecentDocs, OpenSave and LastVisited dialog history, typed paths, Explorer search terms, and Shellbags, which record folders a user browsed, including folders that no longer exist.

Network

Every network the machine joined, with first and last connection times, gateway MAC addresses and whether it was wired or wireless: a movement history for a laptop.

Time, on every key

Because each key carries a last-written timestamp, the registry is not just a state snapshot . It is a partially ordered record of when that state changed.

The catch this guide is really about

All of that assumes the hive file you are reading is the registry that was actually running. Very often it is not.

Windows does not finish a registry write in one go. It writes what it is about to do into a second file first, changes the hive, and only then marks the hive as settled. Copy a hive off a machine that is switched on and you will usually catch it mid-sentence: the file is complete and readable, but the most recent changes are not in it yet. They are in a file sitting right next to it, called SYSTEM.LOG1 or NTUSER.DAT.LOG1, which most tools do not open at all.

The practical result is that the newest registry activity, which is usually the activity closest to whatever you are investigating, is the part most likely to be missing. That is sections 15 to 17.

4. The Registry Is a Set of Files

Here is the step that turns the registry from an abstraction into evidence: the tree you see in regedit does not exist on disk. It is assembled at boot from a handful of separate files, each mounted at a name.

Hive fileLocationMounted atHolds
SYSTEM\Windows\System32\config\HKLM\SYSTEMdrivers, services, devices, mounted volumes
SOFTWARE\Windows\System32\config\HKLM\SOFTWAREinstalled applications, OS configuration
SAM\Windows\System32\config\HKLM\SAMlocal user accounts and group membership
SECURITY\Windows\System32\config\HKLM\SECURITYLSA policy, cached secrets
DEFAULT\Windows\System32\config\HKU\.DEFAULTthe profile in force before anyone logs on
NTUSER.DAT\Users\<user>\HKU\<SID>one per user: settings and activity
UsrClass.dat\Users\<user>\AppData\Local\Microsoft\Windows\HKU\<SID>_Classesshell state, file associations, Shellbags

Two consequences that matter constantly in casework.

  • HKEY_CURRENT_USER is not a file. It is whichever NTUSER.DAT belongs to the account you are running as, mounted under HKEY_USERS. A user who is logged off has no entry there: their hive is on disk, unmounted, and invisible to any live tool.
  • HKLM\SAM and HKLM\SECURITY cannot be read in place even as Administrator; their own ACLs deny it. They must be extracted with backup privileges or read from an image.

So: the tree is a view; the files are the evidence. Everything from here on is about what is inside one of those files.

5. Which File Backs Which HKEY

The previous section said the registry is a set of files. This one is the mapping, because the tree a tool shows you is not laid out like the files underneath it, and three of the branches you can browse are not in any file at all.

What you seeWhat it is on diskNotes
HKLM\SYSTEM%SystemRoot%\System32\config\SYSTEMservices, drivers, device history, mounted devices
HKLM\SOFTWARE%SystemRoot%\System32\config\SOFTWAREinstalled software, most machine-wide persistence
HKLM\SAM%SystemRoot%\System32\config\SAMlocal accounts; the root key denies Administrators
HKLM\SECURITY%SystemRoot%\System32\config\SECURITYLSA secrets, policy
HKU\.DEFAULT%SystemRoot%\System32\config\DEFAULTthe profile the logon screen runs under, not "the default user"
HKU\<SID>C:\Users\<name>\NTUSER.DATone per user, mounted only while that user is logged on
HKU\<SID>_Classes...\AppData\Local\Microsoft\Windows\UsrClass.datper-user file associations, Shellbags
HKCUno filea link to HKU\<SID> of the calling user
HKCRno filea merged view of two Classes subtrees
HKLM\HARDWAREno filerebuilt in memory at every boot

Three things that exist only while Windows is running

HKLM\HARDWARE is volatile. It is built by the kernel at boot from what it finds on the bus, and written to no file. If you are working from an image, it is not missing, it never existed on disk. The same is true of HKLM\SYSTEM\CurrentControlSet, which is a link created at boot to whichever ControlSet00n was selected.

HKCR is a merge, not a hive. It presents HKLM\SOFTWARE\Classes overlaid with the user's UsrClass.dat, with the user's entries winning. A COM hijack that appears in HKCR live may live in either file, and only one of them is per-user. When you work offline you must decide which you are looking at.

WOW6432Node is a redirection, not a location. On a live 64-bit system a 32-bit process asking for HKLM\SOFTWARE\Classes\CLSID is quietly served ...\Classes\Wow6432Node\CLSID. In the file, only the second path exists. A query written against the live illusion returns nothing offline, with no error to say why.

6. The Same Thing, Called Two Different Names

From here on this guide talks about the registry in two vocabularies at once, and it is worth stopping to say so, because almost every confusion about hive files comes from not noticing the switch.

One vocabulary is the one Regedit shows you: keys, subkeys, values, data. The other is what is actually in the file: cells, records, lists, offsets. They describe the same thing. Neither is more correct. But only the second one exists on disk, and a forensic tool reading a hive file has nothing but the second one to work with.

The translation table

Everything in the rest of this guide is one of these six things. When a later section says "the nk record", it means the thing Regedit would draw as a folder.

What Regedit calls itWhat it is in the fileWhat that actually means
The tree under one roota hiveOne file on disk, plus its .LOG files. HKLM\SYSTEM is a file called SYSTEM.
A keyan nk cellOne record holding the key's name, its timestamp, and the positions of everything attached to it. It does not contain its subkeys or values.
A key's subkeysan lh, lf, li or ri listA separate record, elsewhere in the file, holding the positions of the child keys. The parent only knows where that list is.
A valuea vk cellIts name, its type, how long its data is, and where the data lives.
A value's dataa data cell, or inlineJust bytes. Nothing in them says what they mean; the type in the vk record is the only thing that does.
A key's permissionsan sk cellShared. Hundreds of keys with the same permissions all point at one copy.
Where each of those six actually sits
…
The whole file, to scale…

…

The first hive bin, to scale…
Select a row in the table above, or any marker here.

…

The same key, both ways

The table above is the translation. This is the same translation applied to one real key out of the reference system's SYSTEM hive: on the left as Regedit would draw it, on the right as it actually sits in the file. Select any row in either panel and its counterpart lights up. Nothing here needs the byte-level sections that follow; it is here so that when they say nk or lh, you have already seen what they are pointing at.

What Regedit shows, and what is really there
…
What Regedit shows
A tree. Windows assembles it for you.
How it nests, logically
The model in your head. Things inside other things.
What the file holds
Nothing is inside anything. Flat records, joined only by offsets.
Select a row on either side.

Nothing connects directly

The three relations you actually care about are all indirect, and the middle hop is always a record Regedit does not show:

  • A key to its subkey is two hops. The key's subkey list offset names a list, and an entry in that list names the child. The key never mentions the child directly.
  • A key to its value is two hops, in exactly the same shape: the key names a value list, and an entry in that list names the vk.
  • A value to its data is one hop, and sometimes none at all, because a small value keeps its data in the offset field itself.

That is why a key record is only 88 bytes while the key it represents can have thousands of children, and it is why a parser spends most of its time following offsets rather than reading records.

The dimmed rows on the right have no counterpart on the left: they are the subkey lists and the value list. Regedit never shows them and never mentions them, and they are most of the work a parser does. A key does not contain its children; it contains the position of a separate record that lists where the children are.

Two words this guide uses constantly

Both come up on nearly every page from here on, and both are worth pinning down now rather than guessing at.

An offset

A position in the file, counted in bytes from the beginning. Offset 4096 means "4,096 bytes in from the start of the file".

The registry format has no names inside it in the way a filesystem does. When one record needs to refer to another, it stores a number saying where that other record sits. That number is an offset, and following it means jumping to that position in the file and reading what is there.

Nearly every structure in this guide is mostly offsets. A key record is largely a small set of numbers saying where its name, its subkey list, its values and its permissions can be found.

A parser

Any program that reads the hive file directly, byte by byte, instead of asking Windows to read the registry for it.

Regedit is not a parser in this sense. It asks Windows, and Windows hands back a tidy tree. Every forensic tool is a parser, because the machine it is examining is usually switched off, or is not the machine the analyst is sitting at. Crow-Eye is a parser. So is every tool this guide mentions.

This matters because a parser has to make every decision Windows would have made for it, and the sections that follow are largely about the decisions that are easy to get wrong without any error appearing.

Why the difference is the whole point

Regedit shows you a tidy tree because Windows builds one for it. On disk there is no tree: there are records scattered through the file, pointing at each other by position. The tree is something a program assembles by following those pointers.

That is why deleted keys can be recovered, why a half-written hive can look perfectly healthy, and why two tools can read the same file and disagree. All three are covered later, and all three come from the same fact: what Regedit shows you is a reconstruction, not a picture of the file.

7. The Base Block

A hex editor shows a file as what it really is: a long row of numbered bytes, with no structure except the one you know to look for. Open a hive in one and the very first four bytes read regf. That is the format's signature, and checking it is the first thing any program reading a hive does, because if those four bytes are anything else the file is not a hive and nothing that follows will mean what you think it means.

The next 4096 bytes are the base block, which is this format's word for the file header. It is a fixed-size block of fields at fixed positions, and it answers four questions: which version of the format this is, whereabouts in the file the root key sits, how much of the file is actually in use, and, the one that matters most later, whether the file was closed cleanly or was still being written when it was captured.

Click any byte below to see what it is. This is a real SYSTEM hive from a running Windows 11 machine.

Base block: 0x00–0xAF, plus the checksum at 0x1FC
Real bytes from a running machine. The two red fields at the top are the subject of sections 15 to 18.
Field Inspector

The checksum, and why it bites

The last segment of the map is the checksum at 0x1FC: d4 41 93 6d, or 0x6D9341D4 in this hive. It is an XOR of the 127 32-bit words before it. That phrase is doing a lot of work, so unpacking it: everything from the start of the file up to the checksum is 508 bytes, and reading those 508 bytes four at a time gives 127 numbers, each 32 bits wide. XOR them all together and you should get the value stored here.

Why XOR and not a real checksum. XOR is about as cheap as an operation gets, and this value is recomputed every time the header is flushed to disk, which happens constantly. It is there to catch a storage fault: a bit that flipped on the way to the platter, a sector that came back wrong. It is not there to catch anyone. Two bits flipped in the same position of two different words cancel out exactly, so anybody editing a hive deliberately can keep the checksum valid without effort. Treat it as a signal that the header is undamaged, never as a signal that it is unaltered.

Folding 508 bytes into one number
…

Two results are reserved: if the XOR comes out 0 it is stored as 1, and if it comes out 0xFFFFFFFF it is stored as 0xFFFFFFFE. Any tool that writes a hive back, including the log recovery in section 18, must recompute it, or every reader rejects the file. Note also what the map makes obvious: everything from 0xB0 to 0x1FB is reserved and empty, so the header is 4096 bytes of which fewer than two hundred carry anything.

Two timestamps, and only one you can lean on

The last-written FILETIME at 0x0C reads zero in this hive. Across all seven hives on the reference system it is zero on six, and populated on UsrClass.dat, which records 2026-02-05 17:27:47. So it is written sometimes and not others. Check it; never assume it.

The one populated everywhere is last-reorganised, at 0xA8: when Windows last compacted the hive, rewriting free space and reordering cells. On the reference system all seven land within the same fifteen minutes of 2026-08-12: what a single scheduled maintenance pass looks like from the inside.

Neither is the timestamp forensics leans on. That is the per-key one in section 9, maintained on every key, and the closest thing the registry has to a clock.

What that path field really holds

0x30 is easy to read as a label and move on. Put every hive on one machine side by side and it is plainly something else: the file's own path, and on a user hive that includes the account name:

The 0x30 field across seven hives
And the two timestamps beside it, for comparison.
Hive0x30 stored pathlen0x0C written0xA8 reorganised

Look at the lengths. SAM fits in exactly 31 characters; SOFTWARE, SECURITY and DEFAULT are longer, and every one of them has lost its beginning . The field keeps the last 31 characters and drops the rest. That is why SOFTWARE reads emRoot\System32\Config\SOFTWARE, missing its leading backslash and the first four letters of "System".

Why this is worth knowing

A hive carries a record of where it came from. An NTUSER.DAT pulled out of unallocated space, or renamed, or handed to you in a folder called evidence\hive3.dat, still says \??\C:\Users\Admin\ntuser.dat in its own header: naming the account it belonged to, without any external context.

Treat it as a strong lead rather than proof: it is written when the hive is created, so a profile renamed afterwards keeps the old path, and nothing prevents the field being edited. But it is checked far less often than it deserves.

8. Hive Bins and Cells

After the base block, the rest of the file is one big memory allocator. Windows maps the hive into memory and manages space inside it the same way a heap manages RAM, which is exactly why the structures look the way they do.

The file after offset 0x1000 is divided into hive bins (hbin), each a multiple of 4096 bytes. Every bin begins with its own header: the signature hbin, its offset relative to the start of the bins area, and its size.

How a hive file is laid out
Not to scale: a real SYSTEM hive has thousands of bins.
0x0000 Base block 4096 bytes · the header from section 7
0x1000 hive bin 4096 bytes
0x2000 hive bin 4096 bytes
0x3000 hive bin … and so on
inside that first bin
hbin 32-byte bin header
−88 allocated: the root key
−144 allocated: its subkey list
+56 free: reusable space
−32 allocated: a value

Cells sit end to end with no gaps and no index. You find the next one by reading the current one's size and stepping that far, which is why a single wrong size field makes the rest of the bin unreadable.

Reading one real size field
The root cell at file offset 0x1020.
the bytesa8 ff ff ff
little-endian0xFFFFFFA8
as signed 32-bit−88
meaning88 bytes, in use
usable84 bytes

The same four bytes carry two completely different numbers depending on how you read them, and both readings are legitimate. Read a8 ff ff ff as a signed number and it is −88. Read the identical bytes as an unsigned number and it is 4,294,967,208.

Both are true, because a signed 32-bit number uses its top bit to mean "negative" while an unsigned one treats that bit as just another digit, worth two billion. Nothing in the bytes says which convention applies. The format says it, and the program has to know.

This is why getting it wrong is dangerous rather than merely wrong. A program that reads sizes as unsigned does not crash and does not complain. It reads this cell as four gigabytes long, adds four gigabytes to its position to find the next cell, lands far past the end of the file, and stops. It will report having read the hive successfully and will have read one cell of it.

The same four bytes, read two ways
…

Inside a bin sit cells: the unit of allocation, laid end to end. A cell's first four bytes are a signed 32-bit size, and the sign is doing real work:

How you find the next cell, and how that breaks
There is no index. Each cell states its own length, and that is the whole mechanism.

Negative size = allocated. Positive size = free.

The root cell of this hive sits at file offset 0x1020 and its size field reads -88: 88 bytes, in use. Those four bytes are the reason the file can be read at all. There is no index of cells anywhere in a hive, and no table of contents. Every cell simply begins by stating its own length, so a program finds the next cell by reading the current one's size and stepping exactly that far forward. That is the entire navigation mechanism, and it is why one wrong size field makes everything after it in that bin unreadable: the program is no longer landing on cell boundaries.

On top of carrying the length, the sign carries the status. Negative means the cell is in use, positive means it is free. To work out how much space a cell actually gives you, take the absolute value and subtract the four bytes the size field itself occupies, so this 88-byte cell holds 84 bytes of record.

Freeing a cell therefore costs almost nothing: Windows flips the sign and moves on. It does not erase the contents, it does not shuffle anything up to close the gap, and it does not touch the bytes at all.

Two things follow from this design, and both matter forensically:

Hive structure explorer
Walk the real chain: base block → bin → cell → key → subkey list → value.
Structure Detail

9. Key Nodes: the nk Record

A registry key is a cell whose contents start with the two letters nk. Everything the tree shows you about a key is in this one record: its name, its timestamp, how many subkeys and values it has, and where to find them.

OffsetSizeFieldWhat it means
0x002Signaturenk: this cell is a key node
0x022Flagsroot key, symlink, name is ASCII, and so on. This hive's root reads 0x2C
0x048Last writtenFILETIME, and this one is real. The timestamp forensics actually uses
0x104Parent offsetthe cell of the key above this one
0x144Subkey countstable subkeys: 17 for this hive's root
0x184Volatile subkey countsubkeys that exist only in memory and are never written to the file
0x1C4Subkey list offsetthe cell holding the list of children
0x244Value counthow many values this key holds
0x284Value list offsetthe cell holding an array of vk offsets
0x2C4Security offsetthe sk cell with this key's security descriptor
0x482Name lengthbytes of name that follow
0x4CnKey namenot null-terminated: the length field is the only delimiter
nk record: the root key cell, 88 bytes
Real bytes from 0x1020. The cell’s own 4-byte size comes first, then the record inside it.
Field Inspector

A trap worth naming

The subkey list offset is at 0x1C, not 0x18. 0x18 is the volatile subkey count. Reading the wrong one gets you a small integer where an offset should be, which then points somewhere meaningless inside the first bin, and a parser that does not validate will happily read garbage and report it as registry data. This exact mistake was made while writing this guide, and it was caught only because the resulting "subkey list" had a cell size of 7 235 938 bytes and a signature of \x00\x00. Validate signatures; do not trust offsets.

Volatile keys deserve a note of their own. Some registry keys exist only in memory and are never written to any file. HKLM\SYSTEM\CurrentControlSet itself is one, as is Control\hivelist. This is why a live registry and an offline hive legitimately differ: certain keys simply cannot exist in a file, no matter how the image was acquired.

10. Value Records: the vk Record

If nk is a folder, vk is a file. It holds a value's name, its type, how big its data is, and where the data lives, with one space optimisation that catches out anyone writing a parser.

OffsetSizeFieldWhat it means
0x002Signaturevk
0x022Name length0 means the key's unnamed (Default) value
0x044Data sizebottom 31 bits are the length; the top bit changes everything
0x084Data offsetthe cell holding the data: or the data itself
0x0C4Data type1 = REG_SZ, 3 = REG_BINARY, 4 = REG_DWORD, 7 = REG_MULTI_SZ, …
0x102Flagsbit 0 set = the name is ASCII rather than UTF-16
0x14nValue nameagain, length-delimited, not null-terminated

The inline-data trick

Normally a value's vk record does not contain its data. It contains a number saying where the data is, and the program reads that number and jumps to that position in the file to fetch it. That jump is what "following an offset" means, and it happens millions of times in a registry parse.

For very small values that is wasteful. Allocating a whole separate cell to hold four bytes costs more in bookkeeping than the four bytes are worth. So the format has a shortcut: if the top bit of the data-size field is set, the data is not somewhere else at all. It is sitting in the four bytes that would otherwise have held the position. The field that normally says "the data is over there" instead says "the data is right here, and it is this".

The trap is that nothing about the record looks different. A program that always treats that field as a position will, for a REG_DWORD whose value happens to be 1, jump to position 1 in the file, read whatever bytes are sitting in the header there, and report them as the value's data. The length will be right. The type will be right. The data will be nonsense, and no error will be raised anywhere, because from the program's point of view it did exactly what it was told.

The fix is one line and one branch: strip the top bit off before using the length, and use it to decide whether to read the field as data or as a position.

vk record: a real value cell, 32 bytes
The value TypeID at 0x16A0, cell wrapper included.
Field Inspector
Two real values, and what the same field means in each
…

The explorer in section 8 walks to a real example: the value TypeID under ActivationBroker\Plugins, 78 bytes of REG_SZ data, stored in its own cell because it is far too big to inline.

11. Value Types, and Data Too Big to Fit

The vk record carries a type and a length. Both have a trap in them. The documented set means the twelve REG_* type constants that Microsoft publishes and the format specification records, and this field is not limited to them. The length field, meanwhile, has a ceiling above which the data stops being in one piece.

The documented types, and how often they are actually used

Counted across the three hives on the reference system, every allocated vk record, … values in total:

IDNameWhat the bytes areCount

… REG_SZ and REG_DWORD between them account for most of everything.

Types outside the twelve documented ones

… values on the reference system carry a type that is not one of the twelve. They are not corruption, and they are not spread at random: every one of them sits under a device or driver key, in Enum, DeviceClasses, DriverDatabase\DriverPackages or Setup\Upgrade\PnP.

The data itself says what they are. A value typed 0xFFFF0012 holds UTF-16 text; 0xFFFF0007 is always exactly four bytes; 0xFFFF0011 is one byte; 0xFFFF000D is sixteen, the size of a GUID; and 0xFFFF0010 is eight bytes that decode as a FILETIME. That is the device property type space, which Plug and Play stores in the same field.

What it means for a parser: a switch on the type field with twelve cases and no default will silently drop thousands of values on a normal Windows install, and every one of them is device history.

When the data does not fit

A value's data normally lives in one cell. Past a certain size it cannot, and Windows stores a db record instead: a small cell holding a count and a pointer to a list of segment offsets, with the real data split across those segments in order.

The boundary is easier to measure than to look up. Across these three hives the largest value still held in a single cell is … bytes, and the smallest one that needed a db record is …. The real limit sits between the two, which is what you would expect: a cell cannot cross a hive bin boundary, and a bin is 4096 bytes times some whole number.

One real value, split across segments
…

…

The reason for the split is the allocator from section 8. A cell lives inside one hive bin and may not straddle the boundary into the next, so no single cell can be larger than a bin, and a value bigger than that has to be broken up. The db record is how the format keeps the vk record's shape unchanged while the data behind it becomes a list.

What this means for a parser. The vk record looks exactly the same either way. Same signature, same length field, same offset field. The only thing that distinguishes a value with 200 KB of data from one with 20 bytes is what is sitting at the far end of the offset, so the check has to be made every time: read the two bytes at the target and see whether they are db. Skip that check and every large value on the system returns the same wrong thing, silently, and it will be exactly the right length.

These hives hold … of them. A parser that reads the data offset and takes the next length bytes returns the first segment followed by whatever happened to be allocated next, which is a plausible-looking string of exactly the right length and the wrong content.

12. Subkey Lists

An nk record does not contain its children. It contains the offset of a list of them, and there are four kinds of list, because the registry has to stay fast when a key has tens of thousands of subkeys.

SignatureNameEntryPurpose
lfFast leafoffset + first 4 name charactersearly format; compare the hint before following the offset
lhHash leafoffset + name hashthe modern form: this hive's root uses it
liIndex leafoffset onlyno hint; used where hints do not help
riIndex rootoffsets of other listsa list of lists, for keys with more children than one leaf holds

The hint stored beside each offset is the point. To find a subkey by name, Windows compares the cheap hint first and only follows the offset when the hint matches, because following it means a pointer chase into a different part of the mapped file. Entries are kept sorted, so lookup is a binary search rather than a scan.

A real ri: a list of lists
…
…
▾

…

For anyone writing a parser: lf and lh have 8-byte entries, li has 4-byte entries, and ri entries point at further lists that must be walked recursively. Assuming an 8-byte stride for all of them reads every other entry as garbage on an li list, and produces a subkey count that looks plausible.

lh subkey list: the root’s 17 children, 144 bytes
At 0x12C8. After the header, every 8 bytes is one child: four of offset, four of name hash.
Field Inspector

This is the row the map in section 6 dims out. A key does not contain its children: it contains the position of one of these lists, and the list contains the positions of the children. Regedit shows you neither.

All of it, inside each other

Each map above shows one structure flat. What they cannot show is that every one of them lives inside another, and that the way you get from one to the next is by following an offset out of the record you are looking at. That is what a parser does, and it is what this walks.

The nested walk
Every step is real. Click a green pointer field to follow it.
Field Inspector

Two things worth noticing as you walk it.

  • The cell wrapper is always there. Every record begins with four bytes that are not part of the record. They belong to the allocator. "Cell" and "nk" are not two words for the same thing; one contains the other.
  • Nothing is nested in the file. A key's children are not stored inside it; they are somewhere else entirely and it holds their offset. The hierarchy you see in regedit is reconstructed by chasing those offsets, which is exactly why a single broken one makes a hive unreadable.

13. Security Descriptors and Class Names

Every pointer out of an nk record has now been followed except one. The security offset leads to an sk record, and it is the reason a registry with hundreds of thousands of keys is not enormous.

One descriptor, shared by hundreds of keys

An sk cell holds a Windows security descriptor: owner, group, and the ACL that decides who can read or write the key. What makes it interesting is that keys do not get one each.

Consider the arithmetic if they did. This SYSTEM hive has 69,624 keys, and a modest security descriptor runs to a few hundred bytes. One descriptor per key would be tens of megabytes of almost entirely duplicated data in a file that is 23 MB in total. So the format does the obvious thing: identical descriptors are stored once and shared, each carrying a count of how many keys are using it, and the records are threaded together so Windows can walk them when it needs to find an existing match rather than allocate a new one.

HiveKeyssk recordsKeys per descriptorKeys with a class name
The sk ring
…
…

…

Why an examiner should care

A key whose permissions differ from its siblings has its own sk record, and that is visible structurally without reading a single ACL. Where thousands of keys share one descriptor, an outlier is worth a look: weakened permissions on a persistence key is a way to let a non-administrator write to it, and it leaves no trace anywhere else.

Note also what is not here. The sk record stores SIDs, not names. Resolving them against the machine you are sitting at is how an offline account gets confidently mislabelled as a local one.

Class names, the field almost nothing shows you

An nk record can carry a class name, a second string separate from the key's name and stored in its own cell. Most keys have none. On the reference system … do, and in one place the field is not a label at all: it is where the data lives.

The whole of SYSTEM has just … keys with a class name, and they are these:

KeyClass name length

Four of those five are the boot key

JD, Skew1, GBG and Data under ControlSet001\Control\Lsa each hold eight hexadecimal characters in their class name, and nothing in their values. Concatenated in a fixed order and then permuted, those thirty-two characters are the SysKey, the machine's boot key, which is the input to decrypting the password hashes in the SAM hive.

This is the reason the field is worth knowing about. Windows put a secret somewhere that most registry viewers do not render, and a tool that reads names and values and ignores class names cannot see it at all, while the bytes sit in plain text in the file.

The values are not printed on this page. Every other figure here comes from a real machine and is shown; these four are the exception, because together they are that machine's boot key. The lengths are shown instead, and they are the part that makes the point.

Elsewhere the field is mundane: 535 keys in this NTUSER.DAT carry the class Shell, and a thousand Classes\CLSID keys in SOFTWARE carry REG_SZ. Mundane or not, it is a place data can be and most tools will not look.

14. Carving Deleted Keys and Values

The allocator section made the claim that deleted registry data survives. This section is the number behind it, because a claim like that is worth very little without one.

Deleting a key does not erase it. Windows flips its cell's size field from negative to positive, marks it free, and moves on. The signature, the name, the timestamp and the pointers are all still there, and stay there until something else allocates over that space.

That single fact opens a second way to read a hive. Every tool you have used walks the tree: start at the root, follow subkey lists down, and report what you reach. A key that has been unlinked from the tree is invisible to all of them, and it is still in the file. The other way is to ignore the tree entirely and walk the allocator: step through the bins, read each cell's size field, and look at everything the tree can no longer reach. That walk is what the panel below does, one cell at a time.

Every cell in three hives, counted
Walked bin by bin, allocated and free alike.
HiveBinsAllocated cellsFree cellsFree spaceStill an nkStill a vk

…

Walking one real bin, cell by cell
…
Press Walk the bin. Each step reads one size field and decides, from its sign alone, whether that cell is in use or free.

What comes back out

These are real keys, read out of free space in the reference system's SYSTEM hive. Nothing in the tree points at any of them:

CellRecovered key nameLast writtenValues

A deleted key keeps its own last-written timestamp, which is often the more useful half: it dates the activity, not the deletion.

What a carved key is and is not

It is not proof of deletion by a person. Windows frees registry cells constantly by itself: uninstallers, driver updates, profile maintenance and the periodic reorganisation from section 7 all leave freed cells behind. A key in free space means it was removed, not that anyone removed it deliberately.

Its parent is a pointer, not a path. A recovered nk holds its parent's offset, and that cell may itself have been reused. Where the chain breaks, the key's full path is unknown, and inventing one is how a carved artefact turns into a wrong conclusion. Following those pointers is still worth doing, because it works more often than not: in the reference system's SOFTWARE hive, 941 of the 1,451 carved keys walk their parents all the way back to a key that is still live and get a full path out of it. The other 510 keep their name, their timestamp and their values and are recorded with no path at all, because a partial path reads exactly like a real one.

Free does not mean intact. The counts above are cells whose signature still reads nk or vk. Part of a record can be overwritten while its first two bytes survive, so every field has to be sanity-checked before it is believed: a name length of 60,000 or a timestamp in the year 30,000 means you are reading someone else's data.

Nothing in the record says it was deleted. A carved nk is byte for byte the same shape as a live one; the entire claim rests on where the reader found it. An allocator walk that miscounts a single cell size, or a pointer followed back into space that is still in use, produces a table of deleted keys that were never deleted, and no field anywhere in that output would look wrong. The claim has to be checked from outside itself: walk the tree separately, and require every carved offset to be absent from it.

The reorganisation timestamp matters here too. Compacting a hive rewrites its free space, so a hive that was reorganised recently has less recoverable history than one that was not, and the field at 0xA8 tells you which you are holding before you spend an afternoon on it.

15. Why a File Like This Needs a Journal

That is enough structure to see the problem. Everything above is offsets pointing at other offsets, and a single logical change touches several of them at once.

This is the hinge of the guide. Everything before it describes what a hive looks like when it is intact. Everything after it exists because a hive on disk is often not intact, and the parts that say so are the parts most tools skip.

Take something trivial: an application writes a registry value. Follow it through the structures from sections 8 to 12:

  1. The new data does not fit the existing cell, so a new cell is allocated somewhere else in a bin.
  2. The vk record's data offset is updated to point at it.
  3. The old cell's size is flipped positive to free it.
  4. The parent nk's last-written timestamp changes.
  5. If a bin had no room, a new bin is appended and the base block's bins-size updated.

Those are separate writes to separate pages of the file, and they only make sense together. Interrupt them, whether by power loss, a forced reset or a snapshot taken at the wrong instant, and some pages reached the disk while others did not.

The important thing is what that failure looks like from the outside, which is: fine. A hive is not a text file, where a truncated write leaves an obviously ragged ending. Every structure in it is self-describing and independently valid. A vk record whose data offset points at a cell that was never written still has a correct signature, a plausible name length and a type in range. A parser reads it, follows the offset, and returns whatever bytes happen to be sitting at that address. It reports a value. It does not report an error, because from where it is standing nothing is wrong.

Nor does the base block help. Its checksum covers the first 508 bytes of the header and nothing else, so a hive whose body is half-updated still passes the only integrity check the format has. That is the situation the rest of this guide is about: a file that parses cleanly and describes a state the machine was never in.

Two consequences follow, and they are the reason the second half of this guide exists. The first is that Windows needs a way to finish an interrupted write after the fact, which means writing its intentions down somewhere else before it starts. The second is that an examiner needs a way to tell whether the file in front of them is the finished state or the half-finished one, and that needs a flag.

Those five writes, and an interruption
The same five steps listed above, going to disk one at a time.
Adding one registry value is not one write. Press either button to see what the difference costs.

So the registry does what databases and journalling filesystems do: it writes down what it is about to do, somewhere else, first. That somewhere else is the transaction log, and the flag that says whether it is needed is a pair of counters in the base block.

16. The Two Sequence Numbers

You saw them in the base block map at 0x04 and 0x08, coloured red. Here is why there are two of them and not one, and it is worth playing with the control below rather than just reading it.

The write protocol has three steps:

The write protocol
Run it cleanly, or interrupt it and see what the file is left saying.
Sequence1 at 0x04
1888229
Sequence2 at 0x08
1888229
1 · Raise Sequence1, flush the header
2 · Write the changed pages
3 · Match Sequence2, flush again
Equal: the hive was closed cleanly. Nothing to recover.

One comparison, and you know.

Sequence1 == Sequence2 → the hive was closed cleanly.
Sequence1 != Sequence2 → a write was in flight; the log holds the rest.

This is what every tool means when it says "dirty hive". It is not a heuristic . It is two integers and an equality test. A single counter could say how many writes had happened, but never whether the current one had finished.

What it looks like on a machine that is switched on

Seven hives, captured from one Volume Shadow Copy snapshot of a running Windows 11 system:

HiveSequence1Sequence2State
SYSTEM1,888,2301,888,229DIRTY
SOFTWARE5,102,6715,102,670DIRTY
NTUSER.DAT365,687365,686DIRTY
UsrClass.dat251,613251,612DIRTY
DEFAULT197,479197,478DIRTY
SAM382382clean
SECURITY1,0691,069clean

Five of seven are dirty, and every dirty one is dirty by exactly one. That is not five interrupted writes . It is the ordinary steady state of a machine that is running, because there is always another write in flight. The two clean ones are the two nobody writes to: SAM has been written 382 times in the life of this installation, while SYSTEM has been written 1.8 million times.

For an examiner

On a live-acquired image, expect the interesting hives to be dirty. It is the normal condition, not damage, and not evidence of anti-forensics.

17. The Transaction Logs

Next to every hive sit two more files with the same name and the extensions .LOG1 and .LOG2. This is where the outstanding changes live, and where almost every forensic tool stops.

Two formats

The old format, up to Windows 7, marks its dirty data with a DIRT signature. The new format, from Windows 8.1 onwards, is the one described here. A parser must be able to recognise the old one in order to refuse it. Reading a DIRT log with new-format offsets does not raise an error, it reads plausible bytes and writes them into the wrong parts of the hive.

A new-format log file looks like this:

OffsetContents
0a 512-byte copy of the hive's base block: with its own sequence number
512the first log entry (HvLE), then the next, and the next
…unused space, which is not empty. See the stale tail below.

Note the base block copy at the front, and that it carries its own sequence number. That single detail is the crux of the investigation at the end of this section, and the difference between a working recovery and a corrupted hive.

Inside a log entry

HvLE entry header: real bytes
The first entry of a real SYSTEM.LOG1: sequence 1,889,322, 28,160 bytes, 5 dirty pages.
Field Inspector

After the 40-byte header comes a dirty page reference table: one 8-byte pair per page, an offset and a size, and then the page data itself, back to back in the same order. To apply an entry you walk the table and, for each pair, write the next size bytes into the hive at 4096 + offset.

Applying the dirty pages
The five pages this real entry carries, written into the hive.

These are page images, not deltas

Every offset is 4096-aligned and every size is a multiple of 4096. An entry does not say "change these bytes". It says "this page now looks exactly like this". Measured across the 53 entries of one real log: 620 page writes to only 149 distinct offsets, an average of 4.2 rewrites of the same page, with offset 0, the first bin, holding allocation bookkeeping, present in all 53.

So a later entry completely supersedes an earlier one for the same offset. That is why entries must be applied in sequence order, last write wins.

Why a signature is not enough

A log file is not truncated when it is reused. Windows writes entries from offset 512 onward and leaves whatever was there before in place beyond the last one. The tail of a log file is the previous generation's entries: real, structurally valid, correctly hashed entries that are simply old.

The stale tail
SYSTEM.LOG2: 49% of the file is unused, and 9 valid-looking HvLE signatures sit in it.

A parser that scanned for HvLE and applied everything it found would apply 9 stale transactions to SYSTEM and 15 to UsrClass.dat, writing an older generation's pages over a newer hive. It would not crash. It would produce a hive that opens cleanly and is quietly wrong.

Two mechanisms prevent it. The sequence chain: entries run consecutively, and parsing stops at the first entry whose sequence is not the one expected next, so the stale tail is never reached. Marvin32: each entry carries two hashes, one over its body from offset 40 and one over its first 32 bytes, seeded with 0x82EF4D887A4E55C5. Both must reproduce or the entry is not an entry.

If you are implementing this: reproducing those two hashes is the best self-check available. If your Marvin32 and your header offsets are both right, the stored values come out exactly; if either is wrong, they do not. It validates the algorithm and the layout in one step, before you have written a single byte to a hive.

Why there are two logs

Windows alternates between them, so that at any moment one holds a complete, applicable run. Together the pair covers one unbroken sequence range:

FileEntriesRangePages
SYSTEM.LOG1531889322 → 1889374620
SYSTEM.LOG2341889375 → 1889408535

Which file comes first is not fixed

For UsrClass.dat on the same machine it was the other way round: LOG2 covered 251897–251915 and LOG1 continued 251916–251932. Order the pair by the sequence numbers inside them, never by filename. Assuming .LOG1 leads works on most hives and silently produces the wrong result on the rest.

18. A Gap That Should Not Be There

Everything above is the format as it is understood now. This section is how one part of it was actually established, because the documented rule and the real bytes disagreed, and the two possible explanations demanded opposite implementations.

It is worth reading slowly. This is the hardest thing in the guide, and it is here because the answer is less useful than the method: what to do when the documentation and the file cannot both be right.

The recovery rule as commonly described is that log entries continue from where the hive left off. For SYSTEM, with Sequence2 at 1,888,229, the first log entry should have been 1,888,230. It was 1,889,322. Measured from the hive's own position, that first entry sits 1,093 sequences ahead of it. And it was not a one-off:

HiveSequence2First log entryDistance
first entry − Sequence2
SYSTEM1,888,2291,889,3221,093
SOFTWARE5,102,6705,104,5371,867
NTUSER.DAT365,686366,244558
UsrClass.dat251,612251,897285
DEFAULT197,478197,940462

The first suspicion was acquisition: three files copied at three different moments would explain it. They were re-copied from a single snapshot, together. Identical numbers, and the base block checksums validated. This was the real on-disk state.

The gap, drawn to scale
…
Step 1 of 4
Reading A: the transactions are gone

Everything between the hive's position and the log's start has been overwritten. Applying these entries writes page images onto a state they were never computed against. A replay must refuse.

Reading B: the base block lags

The hive's body is further along than its header admits, and the distance is an artefact of measuring from the wrong place. A replay must apply.

Press Next step to walk the evidence.

Reading A: the transactions are gone

Everything between the hive's position and the log's start has been overwritten. The two states do not meet, and applying these entries writes page images onto a state they were never computed against, producing a file that looks valid and is subtly wrong.

→ A replay must refuse.

Reading B: the base block lags

The hive's body is further along than its header admits, the log continues correctly from the real position, and the distance is an artefact of measuring from the wrong place.

→ A replay must apply.

One of these corrupts evidence. There was no way to tell from the file alone which was true.

A hypothesis, tested and discarded

Reading B implies the hive's body is ahead of its base block, which is testable, because the base block declares the hive bins size at 0x28. If the body had grown past what the header records, the file would be bigger than the header claims. It is: SYSTEM holds 23,851,008 bytes of bins where the header says 23,826,432.

Encouraging, until the log entries were checked. The last log entry declares the same bins size as the base block, not the larger figure. If the log's own final view of the hive agrees with the header, the surplus is not unrecorded growth; it is space Windows pre-allocated and has not used. The hypothesis predicted disagreement and found agreement, so it was wrong.

Why guessing was not an option

It is worth being precise about the stakes, because they are asymmetric and neither error announces itself.

Implement Reading A and you refuse to replay logs that were perfectly good. The tool reports the hive as it was found, the analyst never learns that more recent registry activity existed, and nothing anywhere says a decision was made. Evidence is quietly left on the floor.

Implement Reading B and you replay logs that should have been refused. Page images computed against one state are written over a different one, and the output is a hive that parses cleanly, passes its checksum, and contains a mixture of two moments in time. Evidence is quietly manufactured. This is the worse of the two, because the result looks like a successful recovery.

So a coin flip was not available, and neither was picking the reading that was easier to implement.

How it was settled

At that point the specification, the file, and reasoning about it were exhausted. The remaining options were to guess or to find something that already knew, so the implementation was checked against yarp, the reference implementation of registry log recovery, written by the author of the format specification and sharing no code with it.

It recovered all five hives without complaint. That answers the question empirically, and reading how it decides reveals the rule the common description glosses over:

The rule

Log entries chain from the log file's own base-block sequence number, not from the hive's. A log is eligible when its sequence is at or ahead of the hive's Sequence2; only a log older than the hive is refused.

For SYSTEM.LOG1 that number is 1,889,322, exactly its first entry. The log is self-describing. It was never claiming to continue from the hive; it states where it begins.

The gap was never a gap. It was a chain being measured from the wrong end.

Agreement with a reference is not the same as being right, so the bar was set higher: for every hive, the two implementations had to produce a byte-identical recovered file, or refuse for the same reason. They did: five recovered identically, both declining the two clean hives, with every source file's SHA-256 unchanged either side of the run.

What it recovers

HiveEntriesPagesKeys gainedValues gainedLost
SYSTEM871,1550+50
SOFTWARE2503,905+4+40
NTUSER.DAT691,082+10+40
UsrClass.dat36188+1+20
DEFAULT17233000

Ten keys and four values in one user's NTUSER.DAT: the most recent activity there is, and exactly the part of a user hive an examiner cares about most. Without replay it is not "slightly stale"; it is absent, with nothing to indicate anything is missing.

What is still not known

Why does Windows leave the base block hundreds or thousands of sequences behind its log? Reading B is confirmed by behaviour, but the mechanism is still open. The candidates are that the log is reinitialised at a point which does not update the primary's header, or that the primary is flushed on a schedule independent of the base-block sequence. Nothing measured here distinguishes them, and one hypothesis in that family has already been tested and failed.

Settling it would need instrumented hive flushes on a live system, or kernel-side tracing of the registry writer. Neither is available from a snapshot. The recovery is verified; the explanation for one of its inputs is not, and saying so is more useful than a confident guess.

19. What a Hive Supports, and What It Does Not

The structure tells you what changed, never who changed it

Nothing in a hive file records an identity. Not the base block, not a key node, not a value record, not the transaction logs. A sequence number tells you a write happened and in what order; a key's timestamp tells you when the key was last touched. Which account, which process and which session did the touching lives outside the registry entirely.

A recovered cell is data, not a guaranteed record

Carving works because a deleted cell is unlinked rather than erased. That also means the bytes you recover may be a complete value, a partially overwritten one, or the tail of something else that happens to parse. A carved key with a plausible name and an implausible timestamp is usually the second case. Recovery from free space earns a lower confidence than a live key, and it should be reported as such.

Two parsers disagreeing is information, not a bug

A hive with unmerged transaction logs has two legitimate readings: the file as it sits, and the file as Windows would see it after recovery. Tools differ on which one they show, and neither is wrong. When two tools disagree about a registry finding, the logs are usually the reason, and the difference between the readings is itself the evidence of a write that had not yet landed.

A timestamp covers a key, not a value

Key nodes carry a last-write time; value records carry none. Every value under a key shares that one timestamp, and it moves when any of them is written, when a subkey is created, and when a value is deleted. It is an upper bound on everything beneath it — never a time of use for one particular setting.

20. What This Changes in Practice

Collect the logs

.LOG1 and .LOG2 belong in an acquisition beside every hive. A collection that takes only the hive files has discarded the most recent registry activity: the part closest to whatever prompted the investigation.

Say whether it was recovered

Whether a hive was replayed changes what its rows mean. A tool should record, per hive, whether it was dirty, whether logs were present and whether recovery ran, and when it cannot recover, parse anyway and say so rather than fail silently in either direction.

Never write to the evidence

Recovery produces a modified hive by definition. It belongs in a working copy, with the original opened read-only and its hash unchanged before and after.

One snapshot, all the files

A hive and its logs must come from the same acquisition instant. Copy them at three different moments and you will spend a day chasing an inconsistency that you created, as nearly happened here.

The door you come through decides what is in the file

The hives are locked while Windows runs, which is why everything on this page was collected from a shadow copy. It is worth being exact about what the alternative cannot do, because none of it is a matter of privilege. The registry API has no concept of a freed cell, so nothing in section 14 is reachable through it at any privilege level that exists. And some keys deny even an elevated administrator outright: the roots of SAM and SECURITY, and every device Properties subkey. An API walk of those simply returns a smaller tree, with no error anywhere in it. On the reference system an elevated walk of Enum\USB reaches 110 keys and is refused 21 subkeys; read as a file, the same hive gives 868.

So the answer is to stop asking the API and take the file instead, which three routes do: an existing Volume Shadow Copy, raw access to the disk underneath the lock, and a backup-privileged export, where you ask Windows itself to write the key out for you.

Those three are not equivalent, and the difference lands squarely on the two sections above. A shadow copy or a raw read hands you the hive as it sits on disk, allocated cells and freed cells alike. An export is a fresh file: Windows walks the live tree and writes out what it finds, and a freshly written hive has no free space in it at all. Carve an exported hive and you recover nothing, which is a true statement about that file and a false one about the machine it came from. Which door a hive arrived through is therefore part of what it is, and belongs recorded beside it, because it decides what an empty result means.

Reproduce all of it

Every byte on this page came from files any Windows machine has. The hives are locked while Windows runs, so they must be read through a Volume Shadow Copy or raw disk access. From there: the sequence numbers are two uint32 at 0x04 and 0x08; the first hive bin is at 0x1000; the root cell is at 0x1000 + the offset in 0x24; the log's first entry is at offset 512. A hex editor is enough to confirm every table here.

21. Why One Registry Needs Several Parsers

A reasonable question, having got this far: if every artifact in this file is a key with values, why does a forensic tool not have one registry parser? Crow-Eye has four that touch the registry, and the reason is not organisational tidiness. Each one exists because a different part of the problem is genuinely different.

The first split is the door, not the data

The same key can be reached two ways, and they are not the same operation.

RouteDoorSeesCannot see
LiveWindows APIthe merged tree, volatile keys, CurrentControlSet, the redirected 32-bit viewanything about the file: sequence numbers, unwritten log entries, deleted cells
Offlinethe hive fileevery byte in this article, including the parts Windows would have hiddenvolatile keys, which were never written to a file at all

Everything in the second half of this page follows from that difference. A live read cannot tell you a hive was dirty, because there is no hive from its point of view. An offline read cannot show you HKLM\SYSTEM\CurrentControlSet, because that key is assembled at boot and never stored; the file has ControlSet001 and a Select\Current value saying which one is in force. Crow-Eye resolves that itself, and on the reference system Select\Current reads 1, so the offline parser reads ControlSet001.

The two doors disagree, quietly

The library behind an offline read reports a key's unnamed value as the literal string (default), lowercase. The Windows API reports it as an empty name. Nothing errors either way, so the same key produces a different row depending on which door it was read through, and any comparison between the two silently never matches. That is the shape of nearly every live-versus-offline disagreement: not a crash, a difference.

The second split is what is inside the value

This is the one that actually forces separate code, and it is the point section 11 was building to: a value's data is just bytes, and nothing in the registry says what they mean. The type field tells you it is REG_BINARY. It does not tell you the blob is a shell item ID list.

Nothing to decode

A REG_SZ holding a service's image path is finished the moment it is read. Most of the registry is like this, which is why one parser covers most of it.

A structure from somewhere else

A Shellbag value is a shell item ID list: the same structure Windows writes inside a .lnk file. Decoding it is not registry work at all, which is why that code is shared with the shortcut parser rather than living in the registry one. Explained in full: the shell item And the shortcut it also lives in

An entire database

ShimCache is one value holding a serialised store with its own header, version signature and record layout. On the reference system that single value is 282,902 bytes. The registry is only the envelope it arrived in. Explained in full: one ShimCache record

A type Windows never documented

Device keys store Plug and Play property types in the same field as the twelve REG_* types. A reader that only knows the twelve refuses them, and one refused value is enough to lose the key it sat in.

So the four parsers divide along the two axes above, not by artifact:

ParserOpensExists becauseCannot be folded into the others
Live registrythe APIthe door is the Windows APIit can see volatile keys nothing else can
Offline hivethe filethe door is the fileit can see hive state, logs and deleted cells nothing else can
Binary valuea blobthe payload is a structure the registry knows nothing aboutit is called by both doors, and by the shortcut parser too
ShimCacheone valuethe payload is a database, not a valueits record layout changes with the Windows version, not with the hive format

Which is the argument for reading a file format at this depth in the first place. A parser is a claim about a file. The only way to know whether the claim holds is to open the file yourself.

22. Sources and Credit

This page describes a file format that Microsoft has never published. It is legible at all because other people did the work of writing it down, and the honest thing is to say whose work.

Maxim Suhanov, the format specification

The Windows registry file format specification at github.com/msuhanov/regf is the documentation this guide is checked against. The base block layout, the cell and record structures, the four subkey list types and the transaction log formats are all described there in detail, and where this page states an offset with confidence rather than hedging, that is why.

It is also unusually honest documentation: it marks what is uncertain as uncertain, which is what made it possible to tell the difference between a rule that was wrong and a rule this page had misread.

yarp, the independent check

Yet Another Registry Parser, by the same author, at github.com/msuhanov/yarp. It was used as an oracle for section 18: an implementation of log recovery sharing no code with the one being tested, so agreement between them meant something.

It is a development-time tool here and nothing more. It is not in requirements.txt, it is not imported by anything that ships, and no part of the product depends on it. Its job was to answer one question and it did.

The hives themselves

Every number, byte map and table on this page was measured from real hives captured from a running Windows 11 machine through Crow-Eye's own volume shadow copy path, alongside their .LOG1 and .LOG2 files. None of it is an illustrative example, and none of it was typed in by hand: a generator reads the hives and emits the data the page renders, and a separate script re-measures every published figure and fails if the page and the files disagree.

That is also the limitation. These are seven hives from one machine. Where a statement is true of these files rather than of all Windows systems, the page tries to say so, and where something was tested and turned out not to hold, it says that too.