92.75% against 28%. On DFIR-Metric, GPT-4.1 answers certification multiple-choice questions at 92.75% and solves practical CTF-style forensics at 28%, with NIST CFTT string-search tasks in between at 38.52% (arXiv 2505.19973). It is the sharpest knowledge-versus-practice split in the security literature and the best single slide anyone in this market has.
It is also, on inspection, a slide about somebody else's business.
The reason is structural and it is the whole page: incident response data is an RL-environment business, not a labelling business. ExCyTIn-Bench is built on 57 Microsoft Sentinel log tables with human-analyst-constructed bipartite alert-entity graphs as ground truth. CTI-REALM runs its attacks on isolated Azure infrastructure, collects through the Azure Monitor Agent and anonymises afterwards. Neither is a corpus somebody annotated. Both required an estate you own and attacks you run inside it. That is operations and engineering, and it is a different company from the one that recruits twenty experts and a rubric.
What the artefact is
Disk, memory and log evidence in; a timeline, a root cause, a scope and a containment decision out.
The 2026 addition is SIABENCH, "an agentic evaluation framework for security incident analysis" with 160 scenarios — 25 deep-analysis workflows and 135 alert-triage tasks — spanning network and memory forensics, malware analysis across binary, code and PDF formats, phishing and phishing-kit analysis, log analysis and false-alert detection, evaluated across 11 open- and closed-weight models (arXiv 2603.06422).
Its grading methodology and its scores could not be established. That is a real hole, because for this sub-market the grading methodology is the commercial question — it decides which of the three tiers below the artefact sits in.
Is there a verifier
Three tiers, and they price completely differently.
| Layer | Verifier | Evidence |
|---|---|---|
| Artefact extraction — which registry key, which timestamp, which string | Deterministic | DFIR-Metric Module III: 500 disk-analysis tasks adapted from the NIST Computer Forensics Tool Testing Program |
| Timeline reconstruction over a synthetic incident | Checkable against injected ground truth | Cyber Defense Benchmark's deterministic campaign simulator; ExCyTIn's alert-entity graphs |
| Containment and escalation decision | Expert judgement only | — |
The top tier is real and free — CFTT is US-government work, so the answer key costs nothing. The bottom tier is a service. The money is in the middle tier, and the middle tier requires running infrastructure, because "checkable against injected ground truth" means you injected the ground truth, which means you built and ran the incident.
That is the difference between this and Detection engineering, where the verifier runs against MIT-licensed telemetry somebody else recorded. Here you have to make the telemetry. Defense scores 4 on exactly that: an estate with a library of reproducible incidents is genuinely hard to copy and needs refreshing as attacker tradecraft moves. It is the most defensible asset in the defensive half of the dossier.
The licence position helps: NIST CFTT is government work, OTRF Security-Datasets and Atomic Red Team are MIT. Real engagement material is unusable — privileged, client-owned, and the good cases are exactly the ones under NDA.
Is anyone buying
Budget scores 3, and the reason for the cap is precise: every buyer found is hiring incident responders for its own SOC, not buying incident-response data.
- Anthropic: Incident Manager, Detection & Response ($290,000–$365,000), posted in both San Francisco and Zürich; plus Incident Management Lead, Data Center Security, and Insider Risk Investigator.
- OpenAI: Security Engineer, Detection and Response ($293K–$385K), plus Security Engineer, Insider Threat Detection & Response.
- Mercor: the live Cybersecurity Expert listing at $80–90/hr explicitly names "incident response and digital forensics" as one of five task-construction tracks (listing).
The Mercor line is the only confirmed data buyer, and its client is undisclosed. Everything else is a headcount plan. Compare Detection engineering, where a named lab has published the product scope it wants the data for. See The labs as buyers.
What the expert costs
Cost scores 2, and the honest reason is not the one usually given.
| Source | Figure |
|---|---|
| ZipRecruiter, Digital Forensics (Jul 2026) | $59,725/yr · $28.71/hr; median $50,700; range $22,500–$139,000 |
| Same page, Senior Incident Response Analyst | $96,618/yr · $46.45/hr |
| Same page, SOC Analyst | $99,157/yr |
| Anthropic FTE comparator | $290,000–$365,000 |
Treat the $59,725 as unusable. That title bucket contains law-enforcement forensic technicians and eDiscovery staff. Anchor on senior IR at $46.45/hr, which gives roughly a 1.8x arbitrage against Mercor's $85–95 — as good as Threat intelligence and far better than Detection engineering's 1.1–1.3x.
"Capital-intensive" is the usual objection and the numbers do not support it. UK AISI's cyber ranges run on Proxmox VE virtual machines orchestrated by an open-source Inspect plugin, provisioned on AWS m8i instances with nested virtualisation — "a cheaper alternative to bare-metal instance types" (inspect_proxmox_sandbox). An m8i.2xlarge is roughly $0.423/hr with a 1,024 GB gp3 volume adding about $0.11/hr — call it $0.53/hr per concurrency slot, or $2–8/hr once a realistic multi-host range is sized. AISI's own campaign of 122 runs across two ranges and seven models over four days is a few thousand dollars of infrastructure at most.
Infrastructure is not the cost driver. Expert build time is — and so is the operational complexity of running an estate at all: snapshotting, reset, pool concurrency (each eval sample occupies one instance exclusively), and network egress control. AISI had to write an incident report about unsanctioned agent behaviour during cyber testing and is now "building fine-grained network controls into our cyber ranges."
For a small operator that changes the calculus in a specific direction: you do not need capital, you need a platform team. The build:solve ratio is the number to plan against — Hack The Box billed UK AISI £39,366 for the "Cooling Tower" ICS range, which took roughly 15 human-hours to solve. At a 10–40x build-to-solve ratio that is 150–600 build hours, an implied blended £66–260/hour. See The oracle decides everything and Government buyers.
Getting to them
Reach scores 3. The credentials exist and the register is searchable; the counts are not published.
GIAC's DFIR track — GCFA, GCFE, GCIH, GREM — sits inside 290,000+ certifications issued across 60+ types, with a public Certification Holder Directory but no per-certification totals. The SANS DFIR Summit runs annually with a Europe edition and publishes no attendance. Blue Team Village (10,000+ Discord) is the general defensive channel and the best single room; DFIR-specific Discord counts could not be established. BLS does not break incident response out of its 192,900 information security analysts. More at Defensive supply.
Where the benchmarks sit
| Benchmark | Scale | Frontier score |
|---|---|---|
| DFIR-Metric | 700 MCQs + 500 NIST CFTT disk tasks + CTF module | GPT-4.1: 92.75% MCQ / 38.52% NIST string search / 28% CTF |
| SIABENCH (Mar 2026) | 160 scenarios (25 deep, 135 triage), 11 models | Could not establish |
| ExCyTIn-Bench | 57 Sentinel log tables | GPT-5 high reasoning: 56.2% |
| AuditBench | 51 scenarios (25 Atomic Red Team, 26 OpTC) | F1 1.00 lab / 0.45 real classification / 0.25 lateral movement |
Two readings. The 92.75-versus-28 spread is the pitch, and proof scores 3 because it is a scoreboard a new entrant can publish against quickly. And AuditBench's 1.00-to-0.45 collapse between lab and real telemetry is the standing warning for anyone who plans to build the estate cheaply: synthetic incidents that are too clean measure nothing. The measurement gap has the full picture.
What would kill it
You are competing with the buyer's own SOC. Anthropic and OpenAI both run large detection-and-response functions with real incidents in them. Their telemetry is free, current and uncontaminated; yours is constructed.
The privileged cases stay privileged. The best training material in this discipline is under NDA forever. What is left is what you build, and what you build is only as good as your threat modelling.
Ops eats the margin. A range that needs a platform team, network egress controls and per-sample exclusive instances is a company with a cost base, not a marketplace. The $2–8/hr slot cost is trivial; the engineer keeping it alive is not.
Room scores 4 — no named vendor is selling IR evaluation data, and the two published environment builders in this space (AISI's own stack, Hack The Box) are a government and a training company respectively.
Where the record is thin
SIABENCH's grading methodology is the most important unknown on this page: the only 2026 IR benchmark found, and the property that decides whether its 160 scenarios are a product or a rubric exercise is unretrievable.
Nothing is published on wall-clock reset time for a multi-host range, which is the number that converts slot cost into throughput. And no lab has been observed buying IR data from anyone — the entire buyer case rests on one Mercor listing with an undisclosed client. Compare The security read and the shape argument in The one shape that survives.