39.0%. That is GPT-5 on threat-actor attribution in AthenaBench, against 92.0% on knowledge questions and 85.4% on severity assessment in the same suite. It is the largest capability gap this research found in any security discipline, and it is exactly the gap you would expect a well-funded data company to attack.
Do not. Or rather: understand first why nobody has, because the reason is structural and it does not go away with money.
Attribution has no verifier and arguably no ground truth. Three documented facts, each independently fatal to the product version of this business:
- Vendor fragmentation. Across 12,723 reports from 1,626 vendors published between 2000 and 2025, average pairwise Jaccard similarity is 1.3%, 88% of vendors are niche players, and comprehensive actor coverage requires 14 or more distinct vendor sources (arXiv 2602.17458).
- Naming inconsistency produces fragmented ground truth: the same actor carries different names across vendors, and reconciliation is itself an expert judgement.
- Ground truth is often unknowable. Attribution is an abductive judgement under uncertainty. The correct output is a confidence-weighted assessment, and grading a confidence-weighted assessment requires an expert.
So this sub-market cannot become a product. It can be a good service. The rest of the page is about how good.
What the artefact is
Reports in; extracted entities, ATT&CK technique mapping, actor attribution with confidence, campaign linkage and mitigation recommendations out. No customer telemetry needed at any point.
Two 2026 datasets define the shape of a good unit. A manually constructed set of 150 English-language CTI reports, each rendered as a STIX 2.1 graph, yields 4,777 STIX entities, 5,817 relationships and 1,273 attack-pattern entities mapped to 269 unique ATT&CK techniques and sub-techniques; 25 reports were independently assessed by two researchers with substantial inter-rater agreement and disagreements adjudicated to a gold standard (arXiv 2607.23312). CTIConnect adds 1,860 expert-verified QA pairs integrating five heterogeneous sources across nine tasks in three categories — entity linking, multi-document synthesis and entity attribution — evaluated on ten frontier models with temporal splits spanning 2008 to 2025 (arXiv 2510.11974).
The most commercially interesting CTI finding of 2026 is not about models. CTIFoundry rebuilt the corpus rather than the model — materialising a deterministic ontology graph over CVE, CWE, CAPEC and ATT&CK with official cross-references as typed traversable edges, plus a span-grounded, alias-resolved report layer — and lifted an identically-harnessed agent by +0.19 to +0.28 overall F1 on CTIConnect, with a small model on the scaffold beating a flagship on flat retrieval at roughly half the tool calls. Its thesis: "this substrate, not model capability, is the bottleneck" (arXiv 2608.18613).
[UNVERIFIED — inference] That is a direct product thesis: the sellable artefact in CTI is a structured, alias-resolved, provenance-carrying corpus, not annotation labour. It is also a warning, because building it is an engineering project rather than a labelling project, and one team has now published the method.
Is there a verifier
Only for the mechanical layer, and the split is unusually clean.
Entity extraction and ATT&CK mapping have an adjudicated gold standard and can be auto-scored: local open-weight models used as judges against that reference reached κ = 0.803, micro-F1 above 92%, and false-positive rates below 5%. The same paper still concludes that "expert validation remains essential."
Attribution, campaign linkage and mitigation have nothing. Not a weak oracle, as in Vulnerability research, where sensitivity is 60% and specificity 45% — nothing at all. There is no build to trigger, no rule to compile, no flag to recover.
Defense scores 3 on that asymmetry. The absence of a verifier is what stops this becoming a product, and it is simultaneously what stops anyone automating you out of it: a judgement no machine can check is a judgement someone keeps paying a human for. The score is not higher because there is no defensible asset — the reports belong to their publishers and the extraction layer is already solved.
Is anyone buying
Budget scores 3, and the reason for the ceiling is specific and slightly perverse.
OpenAI's Technical Threat Investigator, Threat Intel Engineering ($230K–$385K, San Francisco, with a parallel UK role) is scoped to "independently conduct complex, end-to-end investigations into capable threat actors to understand their behavior, infrastructure, emerging techniques, and how AI is integrated into their workflows" (posting). Anthropic runs Safeguards Policy Analyst, Cyber Harms ($190,000–$285,000), a Senior Safeguards Policy Lead for Cyber Harms in DC, and a Safeguards Enforcement Analyst, Cyber Harm ($285,000–$330,000).
Now read the demand shape. The labs want threat intelligence about misuse of their own models — a corpus only they hold, generated by their own abuse telemetry. A CTI data vendor cannot supply it. What a vendor can supply is the general actor-tracking substrate and the trained judgement, which makes the realistic buyer a CTI vendor or a government rather than a frontier lab [UNVERIFIED — judgement]. See The labs as buyers and Commercial buyers.
What the expert costs
Cost scores 3: the cheapest technical labour of the eight, throttled by the slowest throughput.
| Source | Figure |
|---|---|
| ZipRecruiter, Threat Intelligence Analyst (Aug 2026) | $100,058/yr · $48.10/hr; 25th $77,500 · 75th $116,900 · 90th $137,000 |
| Mercor blended security rate | $85–95/hr |
| OpenAI FTE comparator | $230K–$385K |
That is a 1.8x premium over day-job hourly — the largest of the eight, and four to five times better than Detection engineering's 1.1–1.3x. On recruiting economics alone this is the easiest technical cohort in the dossier to assemble. Paying the crowd compares all eight ladders.
Then the ceiling lands. CTI-HAL used two independent annotators on 81 reports over eight weeks to reach Krippendorff's α = 0.70 — roughly five reports per annotator-week. At $48/hr across 40 hours, that is about $385 of labour per annotated report. A 1,000-report corpus is ~$385,000 of annotation and ~200 annotator-weeks.
That number is the business. Twenty annotators working a full year produce on the order of 5,000 reports. There is no version of this sub-market that scales into a large data company; there is a very good version that runs as a high-margin service line beside something else.
Getting to them
Reach scores 2 — the lowest of the eight, and it is an evidential judgement rather than a slight.
No CTI-specific newsletter, forum or Discord with a published member count was found anywhere in this research. Compare Detection engineering, which has three channels publishing exact numbers. What exists instead:
- ORKL, the public CTI library: an API probe returned results at offset 25,000 and none at 30,000 or beyond, putting the corpus at 25,000–30,000 reports
[UNVERIFIED — own measurement, reproduced this session](ORKL API). Every report carries a named author, which makes ORKL a de-facto register of practising CTI analysts — the best channel on this page, and it is a by-product. - GIAC GCTI, with a searchable holder directory but no published count.
- ISACA — 185,000 members across 188 countries and 225 chapters, which overlaps CTI through the risk-assessment population but belongs properly to Governance, risk and compliance.
- FIRST, the SANS CTI Summit, and the vendor research-blog authorship register — Mandiant, Recorded Future, ESET and the rest.
Where the benchmarks sit
| Benchmark | Scale | Frontier score |
|---|---|---|
| AthenaBench | Dynamic CTI benchmark | GPT-5: 39.0% attribution, 32.6% risk mitigation, 92.0% knowledge, 85.4% severity; combined 66.1% |
| CTIBench | — | GPT-4: 52% correct / 86% plausible on attribution |
| CTIConnect | 1,860 expert-verified QA pairs, 9 tasks, 10 models | Bottleneck shifts between retrieval and evidence utilisation; CTIFoundry lifts F1 by +0.19–0.28 by changing the corpus, not the model |
| STIX 2.1 gold set | 150 reports, 4,777 entities, 269 techniques | LLM judge κ = 0.803, micro-F1 > 92% — on extraction only |
Proof scores 2. The gap is enormous and highly quotable, but the one team that showed how to close it published the method, and a benchmark you cannot grade is a benchmark you cannot win publicly. More at The measurement gap.
What would kill it
The licence position, which is the worst in security. ATT&CK is cleanly licensed. The reports are not. Each vendor report is separately copyrighted, and no aggregator found in this research publishes redistribution terms. A derived corpus built on scraped vendor PDFs is not saleable — see The terms of service bite first.
Throughput. Five reports per annotator-week is a hard physical ceiling on revenue, and it does not improve with better tooling because the slow part is the reading.
Someone builds CTIFoundry commercially. The substrate thesis is published, the ontology sources are free, and the engineering is tractable. Whoever ships that owns the only defensible asset on this page.
Where the record is thin
The ORKL corpus size is my own measurement, not a published figure. AthenaBench and CTIBench scores are carried from earlier research rather than re-derived. And nothing at all is known about what anyone pays per annotated CTI report — the $385 above is labour cost, not price. Room scores 3 because no named vendor occupies this space, which is either an opening or a verdict everyone else reached first. Compare The security read.