Eye Describe · Deep Dive

Continuous Forensic Reconstruction, against EDR and traditional DFIR

Three approaches that read different data, at different times, and keep different amounts of it. Detection watches for an alert and discards the rest. Forensics reads everything, once someone asks. CFR is the one of those three that reads deeply, as forensics does, and never stops, as detection does.

Detection · EDR / SIEM
Endpoint Detection and Response, and Security Information and Event Management — the alerting layer.
Is this event bad?
"Score it, alert on it, clear it if it looks legitimate."

Continuous, cheap at fleet scale, and blind to anything that scores as normal administration.

Forensics · DFIR
What happened on this host?
"Image the disk, parse every artifact, write the report."

As deep as evidence allows, but reactive: it starts only once a breach is already suspected.

Reconstruction · CFR
What actually happened, continuously?
"Parse the same artifacts DFIR would, on a schedule, before anyone asks."

The forensic depth of DFIR, run continuously enough to be there when the question arrives.

0%
Attack detections with no malware at all
Early 2025 — nothing for a signature to match. CrowdStrike.
0%
Year-over-year growth in built-in binary abuse
Living-off-the-land: legitimate tools, used maliciously.
0 days
Median dwell time before discovery (10–14)
Mandiant M-Trends — and a median hides the long tail.
0 hours
GDPR breach-notification window
The SEC allows 4 business days to determine materiality.
01 · Paradigm

The paradigm difference is structural, not incremental #

EDR and DFIR are not an unfinished version of CFR. Each was designed to answer a different question, and the design decisions that follow are why one cannot simply add features to become another.

Design axisEDR / SIEMTraditional DFIRCFR
The question it was built to answerIs this entity malicious, right now?What happened on this specific host?What actually happened, on every endpoint, continuously?
Default stance on new activityTrust it, clear it if it scores legitimateSuspend judgment until examined by handTrust nothing; correlate all of it before judging
What triggers the workA rule or model firingA suspected incidentA schedule, independent of any alert
Depth of readSelected telemetry, filtered at collectionFull artifact set, parsed by hand or by toolFull artifact set, parsed on every run
RetentionShort, to control storage cost at fleet scaleOnly for the case at hand, once collectedParsed and compressed history, kept across the fleet
Unit of outputAn alert with a scoreAn examiner's reportA queryable, correlated timeline with traceable evidence
02 · Timing

When each layer is actually looking #

Map an intrusion from initial access to discovery, and ask which layer is engaged at each phase. EDR is a spike at execution, if the technique trips a rule. DFIR is a block that starts after discovery. CFR runs for the whole of that period.

EDR — alerts only, gaps where activity scores as legitimate Traditional DFIR — begins only after a breach is suspected CFR — continuous, every phase, before the question is asked
Initial accessExecutionPersistenceLog clearingLateral move.StagingExfiltration
EDR
alert
log cleared → blind
DFIR
idle until a breach is suspected
begins
CFR
continuous parsing and correlation, every scheduled run
← median dwell 10–14 days →
03 · Artifact depth

What each layer actually reads #

"Deep" is a claim about specific artifacts, not a marketing word. Here is what each layer captures from the forensic artifacts that carry the most evidentiary weight on Windows — and how long each one survives before the operating system overwrites it on its own.

Table 1 — what each layer captures
ArtifactWhat it provesEDRDFIRCFR
$MFTFile creation, modification and deletion at the volume level, even after the file is goneNot readIf not yet reusedContinuous
USN journalA change log for the volume; shows renames and deletions logs never mentionNot readIf not yet wrappedContinuous
Registry (Run keys, services)Persistence mechanisms, installed software, user activitySelected keysIf not overwrittenContinuous
PrefetchProgram execution, run count, first and last run timeNot readIf not yet evictedContinuous
ShimCacheEvidence a binary existed and its path, even if never executedNot readIf not yet evictedContinuous
AmCacheExecution with SHA-1 hash, install path, and first-seen timeNot readUsually intactContinuous
ShellBagsFolders a user browsed, including on removable or deleted volumesNot readUsually intactContinuous
LNK files / Jump ListsFiles opened, and from which removable media or network shareNot readIf not yet rolledContinuous
SRUMPer-application network and resource usage, back roughly 30 daysNot readIf ≤ 30 days oldContinuous
Windows Event LogsLogons, process creation (if auditing is on), service changesStreamed liveIf not yet wrappedContinuous

EDR's job is to stay cheap at fleet scale, so it streams selected telemetry, mainly event-log-adjacent signals, and discards the rest. CFR reads the same artifact set DFIR would, on a recurring schedule, so the history exists before the question does. Traditional DFIR's column is the one that needs the most scrutiny — see below.

03 · Artifact depth — evidence survival

How long DFIR's evidence survives on its own #

The DFIR column above is conditional for a reason: each of these artifacts is overwritten by the operating system on its own schedule, whether or not an attacker touches it. If collection happens after that window closes, the evidence is simply gone.

Table 2 — how long DFIR's evidence survives on its own
ArtifactTypical window before it's overwrittenWhat overwrites it
PrefetchCapped at roughly the last 1,024 executionsNew program launches evict the oldest entries; can cycle in weeks on an active host
ShimCacheCapped at roughly 1,024 entriesHeld in memory during the session, flushed to the registry at shutdown, oldest evicted the same way
USN journalWraps by size, commonly ≈32 MB by defaultOrdinary disk activity, not just the attacker's; days on a busy volume, longer on a quiet one
SRUM≈30 days, by Windows' own default retentionAutomatically pruned by the operating system on a rolling basis
Windows Event LogsWraps by configured max size and event volumeNew events, once the log file fills; often days on a busy server unless forwarded off-box
$MFT (deleted-file entries)No fixed time — depends on disk activityThe entry is marked free and reused the next time something writes to that part of the volume

Why this matters more than "DFIR is slower": Mandiant's own numbers put median dwell time at 10 to 14 days, but that is a median — plenty of intrusions, especially insider staging or a dormant foothold, run for months. If initial access predates discovery by longer than these windows, the DFIR examiner isn't slow to the artifact — the artifact is gone. No amount of skill, budget or urgency recovers a Prefetch entry the OS has already reused for something else. For an aged intrusion, continuous collection is not a convenience; it is what decides whether the evidence still exists at all.

The boundary, stated plainly: CFR is not retroactive. It preserves what happens after it is deployed and parsing on schedule — it cannot resurrect a host's history from before it was installed. An intrusion that predates deployment is invisible to CFR too, for the same physical reason it's invisible to DFIR: the artifact is already gone. The honest claim is narrower than seeing everything: once it's running, this engine's parses happen faster than these artifacts roll over, so a three-month-old event it already captured is still there, even after the live host's own copy of that evidence is not.

04 · Anti-forensics

What happens when the attacker cleans up #

This is where the categories separate in practice. A technique that defeats one layer often leaves a trace in an artifact that layer never reads. CFR's advantage is not any single artifact resisting tampering — it is having independent sources to notice when one goes missing.

Clear Windows Event Logs

wevtutil cl, or a targeted PowerShell log clear
BlindEDR

The stream it depended on is gone. No alert fires for activity that already scored as legitimate.

DegradedDFIR

Loses the log-based timeline, but a skilled examiner falls back to $MFT and Registry timestamps by hand.

ResilientCFR

Prefetch, AmCache and the USN journal already captured the execution independent of the log.

Timestomping

Rewriting $STANDARD_INFORMATION to hide when a file was really dropped
BlindEDR

Was never looking at file timestamps to begin with.

DegradedDFIR

Catchable, but requires an examiner who knows to diff $STANDARD_INFORMATION against $FILE_NAME by hand.

ResilientCFR

A deterministic rule diffs the two timestamp attributes on every parse and flags the mismatch automatically.

Delete Prefetch files

Removing evidence of execution after the fact
BlindEDR

Never read Prefetch, so there is nothing to lose.

DegradedDFIR

If collection happens after deletion, that execution evidence is gone for good.

ResilientCFR

Already parsed and stored on an earlier scheduled run, before the deletion happened.

Living-off-the-land binaries

PowerShell, WMI, PsExec — legitimate tools used maliciously
BlindEDR

Scores as legitimate administration by design; this is the primary gap driving the current dwell-time reversal.

DegradedDFIR

Findable by an examiner who knows to look, but requires suspecting the host first.

ResilientCFR

Correlated by identity and time against every other artifact, so the chain shows even when no single step looks malicious.

AmCache present, Prefetch absent

Deleting one artifact but missing a correlated one
BlindEDR

Reads neither artifact, so the inconsistency is invisible.

DegradedDFIR

Visible only if an examiner happens to check both artifacts for the same binary.

ResilientCFR

An absence rule expects the pair and flags the gap as a finding in its own right, not a null result.

USN journal deletion

fsutil usn deletejournal, removing the volume change log
BlindEDR

Never depended on it.

DegradedDFIR

Loses that source entirely if it happens before collection.

ResilientCFR

Already parsed on the prior scheduled run; the deletion itself becomes a new, suspicious data point.

05 · Workflow

Time to the first useful query #

The same question — "what did this account do across the fleet in the last 30 days?" — takes a different path depending on which layer answers it.

EDR / SIEM

Minutes — if the data was kept
T+0

Query the console for the account across retained telemetry.

Gap

Anything outside the retention window, or filtered at collection, is already gone.

Ceiling: the answer is only as complete as what was streamed and kept, a small fraction of the host's state.

Traditional DFIR

Days per host, from a cold start
T+0

Confirm a breach is suspected; open a case.

T+hrs

Collect: disk image or live triage package per host in scope.

T+day

Parse each artifact class with separate, single-purpose tools.

T+days

Manually correlate findings across hosts and write the report.

CFR

Seconds — the reconstruction already exists
T−n

Every endpoint already parsed and correlated on its last scheduled run, before the question was asked.

T+0

Query the identity across the fleet's pre-built, normalised timeline.

T+0

Every result traces to its source record, so the answer is checkable, not just fast.

06 · Data economics

Why nobody kept this much history before #

The approach was always correct; it was not affordable. Storing raw artifacts across a fleet indefinitely does not scale. The shift is what gets stored: parsed, normalised rows instead of raw binary structures.

Relative storage cost / endpoint / year
EDRfiltered telemetry, short retention
lowest
CFRparsed & compressed full history
moderate — full parsed history, kept
Raw artifactskept everywhere, always
unaffordable at fleet scale

CFR keeps far more than EDR — the whole parsed artifact history across the fleet, not a filtered slice — so it costs more than EDR's short-retention telemetry, yet a fraction of what storing the raw artifacts everywhere would.

Keeping everything meant storing everything

Raw artifact retention at fleet scale, indefinitely, was never affordable. This is why continuous forensic depth stayed a specialist, per-incident exercise for two decades.

The shift is what gets kept

A parsed, normalised, deduplicated case is a fraction of the raw evidence it came from. Store the rows, not the disk image.

EDR chose the other trade-off

To stay affordable at scale, it filters at collection and ages data out quickly — the opposite compromise, made for a different question.

07 · Defensibility

What backs the finding when it's challenged #

"The model scored it 0.87" and "the examiner's word" both eventually meet a defense attorney, a regulator, or a board. This is where the finding has to hold up.

The claim

Every conclusion, whether reached by a rule engine or an AI layer, is a specific statement: this identity did this thing, at this time.

The evidence it must cite

Not a summary of evidence — the exact source record: a timestamp, a file, a registry value, an executed query.

Absence is a finding, not a silence

If the evidence will not support a claim, that is recorded too, along with the coverage that was actually searched.

Nothing silently discarded

Every record lands in a result or in a named bucket that explains why it was set aside — verifiable from the run log, not taken on trust.

Tamper-evident, end to end: each commit to the record is hash-chained, so re-walking the log detects tampering, even to a human-readable field. The same seal chain backs the EYE_Logs forensic record.

EDRTraditional DFIRCFR
What backs a findingA risk scoreThe examiner's report and their testimonySource records, on a hash-chained, re-walkable log
ReproducibilityModel version-dependent; hard to replay exactlyDepends on the examiner's notes and process disciplineThe same rules against the same evidence reproduce the same result
08 · Further advantages

What continuous reconstruction unlocks #

None of these are features EDR or DFIR could add with a patch. Each follows from being continuous, correlated across the fleet, and queryable after the fact — properties neither was ever built around.

01

Retroactive detection against history already held

When a new technique or IOC is disclosed, the usual question is "were we affected in the last 90 days?" EDR only answers that going forward, once a new rule ships. Write the new Wing today and run it against the correlated history already stored — the answer arrives in the time it takes to query, not a new engagement.

Why the others can't: both require the data to still exist. EDR's retention has usually rolled past the disclosure date; DFIR needs a fresh collection, and by then key artifacts may already be gone.

02

Persistence caught before it ever fires

EDR is execution-triggered: it sees a scheduled task, Run key, or WMI subscription mainly when something reacts to it running. Continuous parsing re-reads persistence locations on every scheduled run regardless of execution, so a dormant backdoor shows up the moment it's planted.

Why the others can't: EDR has nothing to score until the mechanism fires; DFIR isn't looking until a breach is already suspected.

03

Proving a negative, fast

The costliest question after a suspected incident is often "do we have to disclose this." GDPR gives 72 hours; the SEC's rule gives 4 business days of materiality determination. A traceable record that can affirmatively show an identity never touched a data store turns a multi-week engagement into a same-day legal input.

Why the others can't: a risk score isn't evidence of absence, and DFIR's "we didn't find it" carries less weight than "we checked, and here's the coverage."

04

Cross-host correlation over time

Lateral movement and insider staging are fleet-scale patterns: the same removable-media serial across five machines, the same service account authenticating somewhere it never has before. Threading identity across the fleet, not just within a host, is what actually answers insider-risk questions.

Why the others can't: DFIR's unit of work is one host at a time; stitching cases together across a fleet is manual, ad hoc examiner effort.

05

No emergency live-imaging disruption

Traditional collection often means pulling a production machine offline to image it — downtime, plus a live-response window in which RAM and process state change during collection. Background parsing on a schedule means the reconstruction already exists without an emergency collection event.

06

Turns junior analysts into senior-level output

DFIR has a well-documented talent shortage; skilled examiners are scarce and expensive. When correlation and evidence-tracing are automated and every claim cites its source, a junior analyst can query a pre-built timeline instead of needing a senior examiner for every case — a talent-leverage argument, not just a speed one.

07

One dataset, many uses

The same continuously-correlated history serves compliance audits, M&A security due diligence, and non-security legal matters — an IP-theft or wrongful-termination dispute needing an employee's activity history. Traditional DFIR needs a fresh engagement for each; here the dataset already exists and just gets queried differently.

One claim this page does not make: broader endpoint coverage than EDR (legacy systems, OT networks EDR agents can't reach) is a plausible advantage of a lighter-weight parser, but it is not asserted here — it would need verifying against the actual agent footprint, not claimed from architecture alone.

09 · AI-accelerated attacks

The threat is getting stealthier, not slower #

Section 02 cited two measured numbers: 81% of attack detections in early 2025 involved no malware at all, and abuse of built-in binaries grew 126% year over year. Generative AI is a plausible accelerant of that same trend — the reasoning below is forward-looking, not a new statistic, and is labeled that way throughout.

What AI plausibly changes for the attacker

  • Cost of uniqueness drops to near zero. A distinct, per-target command sequence for every intrusion means hash and signature-based indicators, built from prior incidents, cover a shrinking share of what's actually used.
  • Blending in becomes cheap to do well. An attacker can have a model study a specific environment's normal admin patterns — naming conventions, common tools, typical hours — and shape actions to resemble that baseline more closely than a human operator would bother to.
  • The result is more living-off-the-land, not less. Fileless execution, native OS tools, no dropped binary, no static signature to catch — already the dominant pattern, with AI lowering the skill and time cost of doing it convincingly.

Why this engine doesn't erode the same way

  • It was never asking "does this look bad" at collection time. A program running, a file accessed, a persistence key written are recorded because they happened, independent of how convincingly the action was crafted to look normal.
  • Correlation surfaces the chain even when every step looks legitimate. An AI attacker can make each action individually unremarkable, but not edit the fact that one account touched six hosts in sequence within 40 minutes at 3am for the first time in its history.
  • Absence rules don't care how convincing the attacker was. If two artifacts that normally travel together diverge — AmCache with no matching Prefetch entry — that gap is a finding regardless of how well-crafted the individual steps were.

What's measured vs. what's reasoned: the 81% and 126% figures, and the dwell-time chart in Section 02, are measured and cited (Mandiant, CrowdStrike). The link to AI acceleration above is reasoned from that trend, not backed by an AI-specific breach statistic. The underlying trend is documented; the link to AI is a forecast built on that trend, and is offered as one.

10 · AI investigation speed

AI investigation at a precision that didn't exist before #

The slowest, most expensive part of any incident isn't detecting it — it's answering "how far did this go, and what did it touch." That sets Time to Mitigation: containment can't responsibly start until scope is known. An AI layer only shortens that if it's reasoning over ground truth, not filtered telemetry — the piece that didn't exist before continuous, correlated collection did.

Scoping an incident today

Hours to days before containment decisions are safe
T+0

A hypothesis forms: "which hosts did this touch?" No system has that answer pre-built.

T+hrs

The EDR console gives partial telemetry — only for hosts still in its retention window, only what it streamed.

T+hrs

Anything beyond that triggers a formal DFIR collection on the hosts in question.

T+day

An examiner manually cross-references artifacts, host by host, to build the picture by hand.

Result: containment decisions get made against a picture that's still filling in as the investigation continues.

Scoping an incident with the Eye

Minutes — the correlated history already exists
T+0

Ask directly: "every host this identity touched in the last 90 days, in order."

T+0

The query runs against history already parsed and correlated — no new collection, nothing to wait for.

T+0

The full movement graph returns with a cited source record at every hop — a checkable trail, not a summary.

T+0

A follow-up — "what data did it access or move" — is answered from the same pre-built correlation.

Result: containment and disclosure decisions get made against a complete picture within minutes.

Why an EDR copilot doesn't get here: several EDR vendors now ship an AI assistant on top of their console. The gap isn't model quality — it's what the model is allowed to see. An assistant layered on filtered, short-retention telemetry either says "I don't have that" when asked for full blast radius, or worse, produces a plausible-sounding answer ungrounded in evidence the platform never kept. The Eye's advantage isn't a better model — it's being handed ground truth other AI investigators structurally cannot reach, because the platform underneath them was never built to retain it. An AI that reasons over incomplete data and states its conclusions with confidence is a liability in exactly the moment — scoping a breach — where a wrong answer costs the most.

11 · Insider risk

Insider risk becomes visible to management, not just the security operations centre #

Insider risk has historically been caught two ways: a rule-based DLP tool that's too noisy to act on, or a lawsuit after the data has already left. A continuous, identity-threaded history makes a third option possible — the pattern only exists in correlation over time, and only an AI given months of grounded history can narrate it.

SignalSeen aloneSeen correlated, with evidence
Large after-hours data transferA big file copy on a Friday — happens constantly, not actionable aloneFirst time this identity has ever accessed that folder, three weeks after a resignation letter, to a personal cloud endpoint never used before
Removable media useUSB activity — common, mostly benignThe device first appeared the same week as the after-hours access pattern, and was never seen on this host before
Browsing a competitor-sensitive folderOne folder open in ShellBags — could be anythingA folder this identity has never opened in its entire on-file history, opened twice, both outside business hours

Why this is new — it needs two things that never existed together: a real behavioral baseline needs months of consistently parsed history per identity across every artifact class — exactly what Section 03 showed traditional collection can't hold onto, since the underlying artifacts roll over before anyone builds a baseline from them. And it needs an AI able to turn that graph into a plain-English narrative grounded in cited evidence, not a black-box risk score. Most bolted-on UEBA tools sit on the same filtered EDR telemetry — the baseline is as shallow as the data feeding it.

Why it's defensible, not just detectable: "Risk score: 0.87" is not something HR or Legal can put in front of an employee or a court. A narrative where every claim cites its source record — the folder, the timestamp, the device — is the same evidentiary standard the rest of this engine holds itself to, applied to a much earlier, much cheaper point in the story: before the data leaves, not after.

The governance line this has to respect: this is a capability, not a mandate to use it unsupervised. It works as a governance tool HR and Legal invoke under existing employee-monitoring policy and applicable law — not as silent, standing surveillance. Positioned that way, it's also the more defensible sell: management gets an evidenced pattern to review, not an automated accusation.

12 · What this cannot do

The limits of continuous reconstruction #

Every argument above is about what survives and for how long. Four things do not follow from it, and a comparison that claimed them would be making the same mistake it accuses detection of making.

It cannot recover what was already gone. The first parse captures whatever the host still held at that moment, and nothing earlier. An intrusion that predates deployment is bounded by the same artifact lifetimes as any other investigation — the advantage starts on the day collection starts, not retroactively.

It inherits the limits of the artifacts it reads. Parsing a record more often does not make the record more truthful. A timestomped $STANDARD_INFORMATION is timestomped in the reconstruction too; a ShimCache entry still does not prove execution; a Prefetch file still does not name a user. Continuous collection improves survival, not inference — which is why every artifact page in this library carries its own limits section.

It is not detection and does not stop anything. Nothing here alerts, blocks or kills a process. The value is entirely retrospective: a better record of an intrusion is not a defence against one in progress. An estate that replaced EDR with this would answer questions faster and prevent nothing.

Cadence sets the resolution. "Continuous" means on a schedule, not per instruction. Anything that happened and rolled over inside the interval between two parses is as gone as it would have been otherwise, so the interval has to be chosen against the shortest-lived artifact that matters rather than against convenience.

In closing

Where this leaves the category #

None of this makes EDR or DFIR obsolete. EDR is still the fastest way to stop a known-bad binary in its tracks, and a trained examiner is still who a court wants to hear from. What CFR changes is what exists before either of them is engaged: a continuous, correlated, evidence-backed history that an EDR alert can be checked against, and that a DFIR examiner can start from instead of a cold disk image.

The four requirements that make something a CFR platform: continuous artifact parsing, deterministic correlation by identity and time, verifiable evidence traced to source records, and complete accounting of every record the engine touches. A product missing any one of the four is doing collection, triage, or detection — not reconstruction.

Sources: Mandiant M-Trends 2011–2026; CrowdStrike Global Threat Report; CISA living-off-the-land guidance; SANS DFIR artifact references.