The deep dossier found what this screen missed: there is a direct incumbent (Loginsoft, at $25–49/hr), the money is one budget shared with the offensive side rather than a thinner separate one, and the market divides on whether the artefact has an automated oracle — not on the colour of the hat. The scores below are argued down in the section at the end of this page; the replacement scoring is at The security read and the one surviving product shape at The one shape that survives.
Defensive security has the largest published capability gap of any domain in this atlas, and it is not close.
Meta and CrowdStrike built CyberSOCEval jointly — 32 contributors, with Lauren Deason of Meta as corresponding author — covering malware analysis (609 QA pairs from real sandbox detonations of ransomware, RATs and infostealers) and threat-intelligence reasoning (588 QA pairs from 45 threat intel reports mapped to MITRE ATT&CK). Frontier models score 23–34% on malware analysis and 43–53% on threat-intelligence reasoning, and the authors state plainly that "current LLMs are far from saturating our evaluations." Reasoning models showed no dramatic advantage (arXiv 2509.20166).
Read the method, not just the scores. The questions were synthetically generated and then manually validated and edited by cybersecurity experts. That is a frontier lab paying security experts to hand-validate a defensive dataset. It is precisely the product, built once, as a paper, by a company that does not sell it.
And — this sentence originally read "nobody sells it" — searching detection-engineering, SOC-triage and IR vendors across funding news, RL-environment directories and threat-intel vendor lists returned no company selling defensive-security expert data or environments to AI labs. That was wrong. The deep dossier found Loginsoft doing exactly this, and the correction is argued at the end of this page. The adjacent names — Dropzone, Prophet, ReliaQuest, Anomali, Exabeam, Elastic, SOC Prime — are product companies that consume expert judgement internally, and that part holds (The defensive vendors).
What the data actually is
The wedge has to be stated plainly, because getting it wrong kills the company in month four.
You cannot buy the logs. You can buy the narration. Malware samples are freely available from MalwareBazaar and VirusTotal. The samples are not the moat. The analyst's reasoning trace over a sample is. So: pay a $72/hr reverse engineer to detonate a public sample in their own sandbox and narrate the triage — what they looked at first, which string made them change hypothesis, why they dismissed the packer, what they would write as the detection rule and what that rule will false-positive on.
Three product lines follow from that:
| Line | Input | What makes it scarce |
|---|---|---|
| Malware triage traces | Public samples (MalwareBazaar, VirusTotal) | The analyst's ordering of hypotheses, not the verdict |
| Detection-rule rationales | Public Sigma / Elastic rules | Why the rule is written this way, and its known evasions |
| Threat-intel reasoning | Published vendor reports, MITRE ATT&CK | Mapping a narrative report to technique IDs with the disagreements preserved |
None of the three touches customer telemetry, which is the only way this niche is legally buildable at all. It is the same construction Law and Accounting, audit and tax use: public artefact, paid judgement.
Proof scores 5. With a 23–34% baseline published by Meta and CrowdStrike and no canonical benchmark, being visibly the best source here takes one good release, not one good year.
Is anyone buying
Real, but thinner than the offensive side, and you should not conflate the two. Budget scored 3 here originally; the deep dossier moved it to 5, for the reasons at the end of this page.
- CyberSOCEval itself is the strongest signal: two large companies spent expert hours building a defensive eval and published the finding that models fail it (arXiv 2509.20166).
- OpenAI posts a Product Manager, Cyber Defense and Blue Team and a Security Researcher, Agentic AI Threats (OpenAI careers).
- Anthropic's Project Glasswing (7 April 2026) launched with 12 partners including AWS, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, NVIDIA and Palo Alto Networks, plus 40+ critical-software maintainers, committing $100M in model credits, $2.5M to Alpha-Omega/OpenSSF and $1.5M to the Apache Software Foundation (Anthropic). Trade press reports Claude Mythos identified 10,000+ software flaws under the programme
[WEAK](Help Net Security). - Anthropic entered the Collegiate Cyber Defense Competition — a defensive competition — alongside its offensive CTF entries, explicitly to obtain human baselines against professional security researchers (Anthropic).
OpenAI's Red Team Specialist — Cyber posts $198K–$320K plus equity with responsibilities that explicitly include "constructing datasets" (OpenAI) — but that is offensive work. The blue-team roles found carry no published band, and no lab posting hires detection engineers or malware analysts to produce training data. Glasswing spends on model credits and foundations, not on expert-hour contracts you could bid for. The demand is directional; treat a named defensive contract as unproven.
Getting the experts
Small, tight, and unusually easy to reach — the strongest practical argument for this niche. Reach scores 4 (not 5, because there is no single national register the way NASBA is for CPAs).
The SANS/GIAC certification register is the filter: GCIA and GCIH broadly, and GREM specifically, which is close to a perfect filter for malware analysts. Then the repositories where practitioners publish their detection logic under their own names — the Sigma rule repo and Elastic detection-rules on GitHub — which function exactly like GitHub review history does for engineers: a public, verifiable credential in the precise skill you are buying.
Also: Detection Engineering Weekly and its Discord, SANS DFIR Summit, BSides, BlueTeamCon, r/blueteamsec, MISP and OpenCTI community instances, VirusTotal and MalwareBazaar contributor communities, and the Volatility and Ghidra plugin ecosystems.
BLS counts 192,900 US information security analysts at a median $62.11/hr, growing 21% (BLS). The subset you want is far smaller. ISC2's 2025 study surveyed 16,029 practitioners and found 59% reporting critical or significant skills needs, up from 44% in 2024, with AI the number-one skills gap at 41% (ISC2) — a workforce that knows it is behind on exactly your subject.
What it costs to run
Cost scores 4. Malware reverse engineers average $150,631/yr (~$72/hr), with the 25th–75th percentile at $141,834–$159,089 and typically 10+ years of experience required (Salary.com). Expect to pay $120–150/hr for evening hours from someone at that level; the scarcity premium is real.
The split matters. Alert triage at L1/L2 is cheap and the expertise is barely differentiated. Malware analysis and detection engineering pay ~$72/hr for genuinely scarce people, and that is the tier worth building on. Capital requirements are otherwise near zero: sandboxes, a few VMs, and samples that cost nothing.
Who is already there
Almost nobody, and the "almost" is doing more work than this page allowed. Room scored 5 originally and is now 3 — see the correction at the end. CyberSOCEval is a paper, not a business. The public benchmark landscape is fragmentary and academic — OpenSec on incident-response agent calibration under adversarial evidence (arXiv 2601.21083), an attack-investigation benchmark (arXiv 2606.10281), Elastic's agentic SOC framework (Elastic), a Microsoft end-to-end SOC benchmark, Sophos's LLM security benchmarking. None is canonical.
Contrast Offensive security, where the benchmark layer has already been captured by Gray Swan AI and Irregular. The defensive side is the same profession with the same conferences and none of the venture-funded incumbents — though the deep dossier found an unfunded one, Loginsoft, which is the correction at the end of this page. The specialist wedge argues the gap is the best available arbitrage; be honest that the question is open.
What would kill it
SOC telemetry is customer property. SIEM logs, EDR traces and case notes belong to the enterprise, are covered by MSSP contracts, and are dense with employee PII and internal network topology. There is no consent path that scales.
Detection rules are trade secrets. A mature SOC's rule library is the product. The best detection engineers are contractually unable to show you their best work.
Incident reports are privileged. They are frequently produced through outside counsel precisely so they are not discoverable — the same privilege wall as Law, with less room to work around it.
Clearances. Government and defence SOCs exclude a slice of the best people from anything with a foreign customer.
If you ever take a shortcut into real telemetry, one breach notification ends the company. The public-sample discipline is not a nice-to-have; it is the business model.
Defense scores 4: malware families, techniques and detections churn continuously, so the dataset decays on its own and must be refreshed — the healthiest possible shape for a data business, and the opposite of a one-off corpus sale.
The first ninety days here
Query the GIAC register for GREM holders and cross-reference contributor names on the Sigma repo. Recruit 25–40 analysts — this pool is small enough that the first cohort is a list of individuals, not a funnel. Pay $130/hr and put "public samples only, no employer data, ever" in the first line of the brief, which will do more for trust in this community than the rate.
Then build the thing that does not exist: a public, rubric-graded SOC benchmark covering triage ordering, rule authoring and intel-to-ATT&CK mapping, sitting alongside CyberSOCEval rather than duplicating it, and reporting where the 23–34% comes from rather than restating it. Publish the disagreements between analysts. See The first ninety days.
Two scores the deep dossier moved, and why
room: 5 → 3. The claim above that "nobody sells it" was a search failure, not a finding. Loginsoft has been selling "Security Data for AI Training" — "curated, labeled, and synthetic cybersecurity datasets" with "expert labeling and ground-truth validation" — plus a separate AI Model Validation line with "human-in-the-loop review" (Loginsoft; model validation). Founded 2005, 201–500 people, Hyderabad development centre, billing $25–49/hour [WEAK — third-party directory] (Enosis). It has no venture capital, no benchmark, no lab logo and names no client for that line — so it is the price floor rather than the market — but a category with a trading incumbent at a third of Mercor's rate does not score 5. Add the six external security research firms Anthropic already pays to triage its disclosure pipeline, and the position is occupied at both the cheap end and the expensive one. See The defensive vendors and What the dossier could not establish.
budget: 3 → 5. This page treated defensive as a separate, thinner market on the ground that the named data-producing req was offensive. That framing was wrong. It is the same two buyers and the same evaluation budget — one of which paused its largest planned frontier RL run on cyber grounds in August 2026, and the other of which published that its models have "saturated all of our current cyber evaluations." Anthropic and OpenAI carry roughly 80 cyber-titled open roles between them, topping out at $485,000–$755,000. The labs as buyers has the full table.
The other four scores survive. proof: 5 holds — the 23–34% gap is still published and still open, though Microsoft and CrowdStrike giving away high-quality defensive benchmarks makes the benchmark route narrower than it looked (The measurement gap). cost: 4 and reach: 4 hold with caveats now documented in Defensive supply. defense: 4 holds because the deep look found the structural reason it should: where there is no automated oracle, the labour never automates away (The one shape that survives).
Where the record is thin
The pool size you actually care about is unmeasured. BLS's 192,900 covers all information security analysts; nobody counts detection engineers or malware reverse engineers separately, and the GREM holder count is not public.
No defensive contract with a lab has been observed at any price, so the payRate above is a salary-survey inference and not an observed clearing price. Salary.com is a self-reported aggregator. And the question that decides the niche is unchanged: no lab publishes what it pays per hour or per artefact for expert cyber work anywhere, which means gross margin cannot be modelled from public information at all. See Sizing the cyber pot.