Three approaches that read different data, at different times, and keep different amounts of it. Detection watches for an alert and discards the rest. Forensics reads everything, once someone asks. CFR is the one of those three that reads deeply, as forensics does, and never stops, as detection does.
"Score it, alert on it, clear it if it looks legitimate."
Continuous, cheap at fleet scale, and blind to anything that scores as normal administration.
"Image the disk, parse every artifact, write the report."
As deep as evidence allows, but reactive: it starts only once a breach is already suspected.
"Parse the same artifacts DFIR would, on a schedule, before anyone asks."
The forensic depth of DFIR, run continuously enough to be there when the question arrives.
EDR and DFIR are not an unfinished version of CFR. Each was designed to answer a different question, and the design decisions that follow are why one cannot simply add features to become another.
| Design axis | EDR / SIEM | Traditional DFIR | CFR |
|---|---|---|---|
| The question it was built to answer | Is this entity malicious, right now? | What happened on this specific host? | What actually happened, on every endpoint, continuously? |
| Default stance on new activity | Trust it, clear it if it scores legitimate | Suspend judgment until examined by hand | Trust nothing; correlate all of it before judging |
| What triggers the work | A rule or model firing | A suspected incident | A schedule, independent of any alert |
| Depth of read | Selected telemetry, filtered at collection | Full artifact set, parsed by hand or by tool | Full artifact set, parsed on every run |
| Retention | Short, to control storage cost at fleet scale | Only for the case at hand, once collected | Parsed and compressed history, kept across the fleet |
| Unit of output | An alert with a score | An examiner's report | A queryable, correlated timeline with traceable evidence |
Map an intrusion from initial access to discovery, and ask which layer is engaged at each phase. EDR is a spike at execution, if the technique trips a rule. DFIR is a block that starts after discovery. CFR runs for the whole of that period.
"Deep" is a claim about specific artifacts, not a marketing word. Here is what each layer captures from the forensic artifacts that carry the most evidentiary weight on Windows — and how long each one survives before the operating system overwrites it on its own.
| Artifact | What it proves | EDR | DFIR | CFR |
|---|---|---|---|---|
| $MFT | File creation, modification and deletion at the volume level, even after the file is gone | Not read | If not yet reused | Continuous |
| USN journal | A change log for the volume; shows renames and deletions logs never mention | Not read | If not yet wrapped | Continuous |
| Registry (Run keys, services) | Persistence mechanisms, installed software, user activity | Selected keys | If not overwritten | Continuous |
| Prefetch | Program execution, run count, first and last run time | Not read | If not yet evicted | Continuous |
| ShimCache | Evidence a binary existed and its path, even if never executed | Not read | If not yet evicted | Continuous |
| AmCache | Execution with SHA-1 hash, install path, and first-seen time | Not read | Usually intact | Continuous |
| ShellBags | Folders a user browsed, including on removable or deleted volumes | Not read | Usually intact | Continuous |
| LNK files / Jump Lists | Files opened, and from which removable media or network share | Not read | If not yet rolled | Continuous |
| SRUM | Per-application network and resource usage, back roughly 30 days | Not read | If ≤ 30 days old | Continuous |
| Windows Event Logs | Logons, process creation (if auditing is on), service changes | Streamed live | If not yet wrapped | Continuous |
EDR's job is to stay cheap at fleet scale, so it streams selected telemetry, mainly event-log-adjacent signals, and discards the rest. CFR reads the same artifact set DFIR would, on a recurring schedule, so the history exists before the question does. Traditional DFIR's column is the one that needs the most scrutiny — see below.
The DFIR column above is conditional for a reason: each of these artifacts is overwritten by the operating system on its own schedule, whether or not an attacker touches it. If collection happens after that window closes, the evidence is simply gone.
| Artifact | Typical window before it's overwritten | What overwrites it |
|---|---|---|
| Prefetch | Capped at roughly the last 1,024 executions | New program launches evict the oldest entries; can cycle in weeks on an active host |
| ShimCache | Capped at roughly 1,024 entries | Held in memory during the session, flushed to the registry at shutdown, oldest evicted the same way |
| USN journal | Wraps by size, commonly ≈32 MB by default | Ordinary disk activity, not just the attacker's; days on a busy volume, longer on a quiet one |
| SRUM | ≈30 days, by Windows' own default retention | Automatically pruned by the operating system on a rolling basis |
| Windows Event Logs | Wraps by configured max size and event volume | New events, once the log file fills; often days on a busy server unless forwarded off-box |
| $MFT (deleted-file entries) | No fixed time — depends on disk activity | The entry is marked free and reused the next time something writes to that part of the volume |
Why this matters more than "DFIR is slower": Mandiant's own numbers put median dwell time at 10 to 14 days, but that is a median — plenty of intrusions, especially insider staging or a dormant foothold, run for months. If initial access predates discovery by longer than these windows, the DFIR examiner isn't slow to the artifact — the artifact is gone. No amount of skill, budget or urgency recovers a Prefetch entry the OS has already reused for something else. For an aged intrusion, continuous collection is not a convenience; it is what decides whether the evidence still exists at all.
The boundary, stated plainly: CFR is not retroactive. It preserves what happens after it is deployed and parsing on schedule — it cannot resurrect a host's history from before it was installed. An intrusion that predates deployment is invisible to CFR too, for the same physical reason it's invisible to DFIR: the artifact is already gone. The honest claim is narrower than seeing everything: once it's running, this engine's parses happen faster than these artifacts roll over, so a three-month-old event it already captured is still there, even after the live host's own copy of that evidence is not.
This is where the categories separate in practice. A technique that defeats one layer often leaves a trace in an artifact that layer never reads. CFR's advantage is not any single artifact resisting tampering — it is having independent sources to notice when one goes missing.
The stream it depended on is gone. No alert fires for activity that already scored as legitimate.
Loses the log-based timeline, but a skilled examiner falls back to $MFT and Registry timestamps by hand.
Prefetch, AmCache and the USN journal already captured the execution independent of the log.
Was never looking at file timestamps to begin with.
Catchable, but requires an examiner who knows to diff $STANDARD_INFORMATION against $FILE_NAME by hand.
A deterministic rule diffs the two timestamp attributes on every parse and flags the mismatch automatically.
Never read Prefetch, so there is nothing to lose.
If collection happens after deletion, that execution evidence is gone for good.
Already parsed and stored on an earlier scheduled run, before the deletion happened.
Scores as legitimate administration by design; this is the primary gap driving the current dwell-time reversal.
Findable by an examiner who knows to look, but requires suspecting the host first.
Correlated by identity and time against every other artifact, so the chain shows even when no single step looks malicious.
Reads neither artifact, so the inconsistency is invisible.
Visible only if an examiner happens to check both artifacts for the same binary.
An absence rule expects the pair and flags the gap as a finding in its own right, not a null result.
Never depended on it.
Loses that source entirely if it happens before collection.
Already parsed on the prior scheduled run; the deletion itself becomes a new, suspicious data point.
The same question — "what did this account do across the fleet in the last 30 days?" — takes a different path depending on which layer answers it.
Query the console for the account across retained telemetry.
Anything outside the retention window, or filtered at collection, is already gone.
Confirm a breach is suspected; open a case.
Collect: disk image or live triage package per host in scope.
Parse each artifact class with separate, single-purpose tools.
Manually correlate findings across hosts and write the report.
Every endpoint already parsed and correlated on its last scheduled run, before the question was asked.
Query the identity across the fleet's pre-built, normalised timeline.
Every result traces to its source record, so the answer is checkable, not just fast.
The approach was always correct; it was not affordable. Storing raw artifacts across a fleet indefinitely does not scale. The shift is what gets stored: parsed, normalised rows instead of raw binary structures.
CFR keeps far more than EDR — the whole parsed artifact history across the fleet, not a filtered slice — so it costs more than EDR's short-retention telemetry, yet a fraction of what storing the raw artifacts everywhere would.
Raw artifact retention at fleet scale, indefinitely, was never affordable. This is why continuous forensic depth stayed a specialist, per-incident exercise for two decades.
A parsed, normalised, deduplicated case is a fraction of the raw evidence it came from. Store the rows, not the disk image.
To stay affordable at scale, it filters at collection and ages data out quickly — the opposite compromise, made for a different question.
"The model scored it 0.87" and "the examiner's word" both eventually meet a defense attorney, a regulator, or a board. This is where the finding has to hold up.
Every conclusion, whether reached by a rule engine or an AI layer, is a specific statement: this identity did this thing, at this time.
Not a summary of evidence — the exact source record: a timestamp, a file, a registry value, an executed query.
If the evidence will not support a claim, that is recorded too, along with the coverage that was actually searched.
Every record lands in a result or in a named bucket that explains why it was set aside — verifiable from the run log, not taken on trust.
Tamper-evident, end to end: each commit to the record is hash-chained, so re-walking the log detects tampering, even to a human-readable field. The same seal chain backs the EYE_Logs forensic record.
| EDR | Traditional DFIR | CFR | |
|---|---|---|---|
| What backs a finding | A risk score | The examiner's report and their testimony | Source records, on a hash-chained, re-walkable log |
| Reproducibility | Model version-dependent; hard to replay exactly | Depends on the examiner's notes and process discipline | The same rules against the same evidence reproduce the same result |
None of these are features EDR or DFIR could add with a patch. Each follows from being continuous, correlated across the fleet, and queryable after the fact — properties neither was ever built around.
When a new technique or IOC is disclosed, the usual question is "were we affected in the last 90 days?" EDR only answers that going forward, once a new rule ships. Write the new Wing today and run it against the correlated history already stored — the answer arrives in the time it takes to query, not a new engagement.
Why the others can't: both require the data to still exist. EDR's retention has usually rolled past the disclosure date; DFIR needs a fresh collection, and by then key artifacts may already be gone.
EDR is execution-triggered: it sees a scheduled task, Run key, or WMI subscription mainly when something reacts to it running. Continuous parsing re-reads persistence locations on every scheduled run regardless of execution, so a dormant backdoor shows up the moment it's planted.
Why the others can't: EDR has nothing to score until the mechanism fires; DFIR isn't looking until a breach is already suspected.
The costliest question after a suspected incident is often "do we have to disclose this." GDPR gives 72 hours; the SEC's rule gives 4 business days of materiality determination. A traceable record that can affirmatively show an identity never touched a data store turns a multi-week engagement into a same-day legal input.
Why the others can't: a risk score isn't evidence of absence, and DFIR's "we didn't find it" carries less weight than "we checked, and here's the coverage."
Lateral movement and insider staging are fleet-scale patterns: the same removable-media serial across five machines, the same service account authenticating somewhere it never has before. Threading identity across the fleet, not just within a host, is what actually answers insider-risk questions.
Why the others can't: DFIR's unit of work is one host at a time; stitching cases together across a fleet is manual, ad hoc examiner effort.
Traditional collection often means pulling a production machine offline to image it — downtime, plus a live-response window in which RAM and process state change during collection. Background parsing on a schedule means the reconstruction already exists without an emergency collection event.
DFIR has a well-documented talent shortage; skilled examiners are scarce and expensive. When correlation and evidence-tracing are automated and every claim cites its source, a junior analyst can query a pre-built timeline instead of needing a senior examiner for every case — a talent-leverage argument, not just a speed one.
The same continuously-correlated history serves compliance audits, M&A security due diligence, and non-security legal matters — an IP-theft or wrongful-termination dispute needing an employee's activity history. Traditional DFIR needs a fresh engagement for each; here the dataset already exists and just gets queried differently.
One claim this page does not make: broader endpoint coverage than EDR (legacy systems, OT networks EDR agents can't reach) is a plausible advantage of a lighter-weight parser, but it is not asserted here — it would need verifying against the actual agent footprint, not claimed from architecture alone.
Section 02 cited two measured numbers: 81% of attack detections in early 2025 involved no malware at all, and abuse of built-in binaries grew 126% year over year. Generative AI is a plausible accelerant of that same trend — the reasoning below is forward-looking, not a new statistic, and is labeled that way throughout.
What's measured vs. what's reasoned: the 81% and 126% figures, and the dwell-time chart in Section 02, are measured and cited (Mandiant, CrowdStrike). The link to AI acceleration above is reasoned from that trend, not backed by an AI-specific breach statistic. The underlying trend is documented; the link to AI is a forecast built on that trend, and is offered as one.
The slowest, most expensive part of any incident isn't detecting it — it's answering "how far did this go, and what did it touch." That sets Time to Mitigation: containment can't responsibly start until scope is known. An AI layer only shortens that if it's reasoning over ground truth, not filtered telemetry — the piece that didn't exist before continuous, correlated collection did.
A hypothesis forms: "which hosts did this touch?" No system has that answer pre-built.
The EDR console gives partial telemetry — only for hosts still in its retention window, only what it streamed.
Anything beyond that triggers a formal DFIR collection on the hosts in question.
An examiner manually cross-references artifacts, host by host, to build the picture by hand.
Ask directly: "every host this identity touched in the last 90 days, in order."
The query runs against history already parsed and correlated — no new collection, nothing to wait for.
The full movement graph returns with a cited source record at every hop — a checkable trail, not a summary.
A follow-up — "what data did it access or move" — is answered from the same pre-built correlation.
Why an EDR copilot doesn't get here: several EDR vendors now ship an AI assistant on top of their console. The gap isn't model quality — it's what the model is allowed to see. An assistant layered on filtered, short-retention telemetry either says "I don't have that" when asked for full blast radius, or worse, produces a plausible-sounding answer ungrounded in evidence the platform never kept. The Eye's advantage isn't a better model — it's being handed ground truth other AI investigators structurally cannot reach, because the platform underneath them was never built to retain it. An AI that reasons over incomplete data and states its conclusions with confidence is a liability in exactly the moment — scoping a breach — where a wrong answer costs the most.
Insider risk has historically been caught two ways: a rule-based DLP tool that's too noisy to act on, or a lawsuit after the data has already left. A continuous, identity-threaded history makes a third option possible — the pattern only exists in correlation over time, and only an AI given months of grounded history can narrate it.
| Signal | Seen alone | Seen correlated, with evidence |
|---|---|---|
| Large after-hours data transfer | A big file copy on a Friday — happens constantly, not actionable alone | First time this identity has ever accessed that folder, three weeks after a resignation letter, to a personal cloud endpoint never used before |
| Removable media use | USB activity — common, mostly benign | The device first appeared the same week as the after-hours access pattern, and was never seen on this host before |
| Browsing a competitor-sensitive folder | One folder open in ShellBags — could be anything | A folder this identity has never opened in its entire on-file history, opened twice, both outside business hours |
Why this is new — it needs two things that never existed together: a real behavioral baseline needs months of consistently parsed history per identity across every artifact class — exactly what Section 03 showed traditional collection can't hold onto, since the underlying artifacts roll over before anyone builds a baseline from them. And it needs an AI able to turn that graph into a plain-English narrative grounded in cited evidence, not a black-box risk score. Most bolted-on UEBA tools sit on the same filtered EDR telemetry — the baseline is as shallow as the data feeding it.
Why it's defensible, not just detectable: "Risk score: 0.87" is not something HR or Legal can put in front of an employee or a court. A narrative where every claim cites its source record — the folder, the timestamp, the device — is the same evidentiary standard the rest of this engine holds itself to, applied to a much earlier, much cheaper point in the story: before the data leaves, not after.
The governance line this has to respect: this is a capability, not a mandate to use it unsupervised. It works as a governance tool HR and Legal invoke under existing employee-monitoring policy and applicable law — not as silent, standing surveillance. Positioned that way, it's also the more defensible sell: management gets an evidenced pattern to review, not an automated accusation.
Every argument above is about what survives and for how long. Four things do not follow from it, and a comparison that claimed them would be making the same mistake it accuses detection of making.
It cannot recover what was already gone. The first parse captures whatever the host still held at that moment, and nothing earlier. An intrusion that predates deployment is bounded by the same artifact lifetimes as any other investigation — the advantage starts on the day collection starts, not retroactively.
It inherits the limits of the artifacts it reads.
Parsing a record more often does not make the record more truthful. A timestomped
$STANDARD_INFORMATION is timestomped in the reconstruction too; a ShimCache entry
still does not prove execution; a Prefetch file still does not name a user. Continuous
collection improves survival, not inference — which is why every
artifact page in this library carries its own limits section.
It is not detection and does not stop anything. Nothing here alerts, blocks or kills a process. The value is entirely retrospective: a better record of an intrusion is not a defence against one in progress. An estate that replaced EDR with this would answer questions faster and prevent nothing.
Cadence sets the resolution. "Continuous" means on a schedule, not per instruction. Anything that happened and rolled over inside the interval between two parses is as gone as it would have been otherwise, so the interval has to be chosen against the shortest-lived artifact that matters rather than against convenience.
None of this makes EDR or DFIR obsolete. EDR is still the fastest way to stop a known-bad binary in its tracks, and a trained examiner is still who a court wants to hear from. What CFR changes is what exists before either of them is engaged: a continuous, correlated, evidence-backed history that an EDR alert can be checked against, and that a DFIR examiner can start from instead of a cold disk image.
The four requirements that make something a CFR platform: continuous artifact parsing, deterministic correlation by identity and time, verifiable evidence traced to source records, and complete accounting of every record the engine touches. A product missing any one of the four is doing collection, triage, or detection — not reconstruction.
Sources: Mandiant M-Trends 2011–2026; CrowdStrike Global Threat Report; CISA living-off-the-land guidance; SANS DFIR artifact references.