No contract value for cyber evaluation exists anywhere in public. Not one lab discloses cyber evaluation spend. Not one evaluation vendor discloses revenue. Not one lab-to-vendor contract value has been published — not Irregular's, not Gray Swan AI's, not Trajectory Labs', not 10a Labs'. OpenAI confirms it pays external testers through "direct payment and/or API credits" and that "[n]o payment is ever contingent on the results," and discloses no amounts (OpenAI).
So everything below is inference from adjacent disclosures. The working is shown step by step so it can be disagreed with at any particular step rather than rejected wholesale.
What is known for certain
| Fact | Value | Source |
|---|---|---|
| AIxCC prize money distributed | $8.5M finals + $1.4M ARPA-H | DARPA |
| Lab in-kind contribution to AIxCC | $350,000 each, Anthropic / Google / OpenAI, in credits | DARPA |
| ARIA adversarial cyber red team contract | £2–3M for ~14 months, one partner | ARIA |
| ARIA total cyber programme | £20M | ARIA |
| Irregular raised / valuation / headcount | $80M / $450M / ~35 people | TechCrunch |
| Gray Swan Series A / valuation / customers | $40M / ~$200M / 20+ | Forbes |
| Cyber expert contributor rate at Mercor | $70–$90/hr vuln analysis; $80–$90/hr senior scenario authoring | Mercor |
| UK AISI range build, human input | ~20 human-hours (corporate) and ~15 (ICS) to solve; built by SpecterOps and Hack The Box | UK AISI |
| Expert cyber range build, academic | 6 senior experts (5+ yrs each) for 266 instances / 8 ranges / 156 hosts | AgentCyberRange |
Everything after this point is derived from that table plus the payroll evidence in The labs as buyers.
Method one: from vendor capitalisation
Irregular and Gray Swan are the two named cyber-relevant evaluation vendors serving frontier labs, carrying $120M of combined venture funding between them, with Irregular at roughly 35 employees and Gray Swan undisclosed but small. Sequoia states Irregular is "already generating millions in revenue" (Sequoia) — a partner blog post, not an audited figure, so directional at best. Frontier labs are "a majority" of Gray Swan's revenue.
Venture-funded vendors at Series A typically run revenue in the single-digit to low-double-digit millions. If the two together do $20–50M of annual revenue and labs are most of it, the directly-contracted third-party cyber and adversarial evaluation market from frontier labs is on the order of $15–40M/yr.
Adding the smaller suppliers named in the Claude Opus 5 card — Trajectory Labs at ~100 hours, 10a Labs at ~16 hours — does not move it. Those are five-figure engagements. [WEAK] — inference from funding stage, not from disclosed revenue.
Method two: from expert labour cost
This is the independent path, and it starts from a real posted rate rather than a valuation.
Cyber experts on Mercor bill $70–$90/hr. Take the UK AISI corporate range as the unit: 32 steps, roughly 20 human-hours to solve. Build effort for an environment of that quality is realistically 10–40x the solve time — you are authoring the network, the misconfigurations, the plausible defensive posture, the scoring, and then verifying it is actually solvable and actually hard. That gives 200–800 expert-hours per premium environment, or $15K–$70K of raw expert labour before infrastructure, validation and margin.
A lab needing ~100 novel environments a year to stay ahead of saturation — a plausible figure given Anthropic retired CyberGym outright between Opus 4.6 and Opus 5 — is looking at $1.5M–$7M of raw expert labour per lab per year, or roughly $4M–$20M at a vendor's gross margin. Across six labs that is $25M–$120M/yr.
The 10–40x build multiplier is the load-bearing assumption and it is mine, not a source's. Halve it and the band halves.
Method three, as a cross-check: from internal headcount
Anthropic and OpenAI carry roughly 80 cyber- and security-titled open roles between them at a rough midpoint of $340K base. Open roles are not filled roles, but assume the filled cyber-relevant population across both labs is 150–300 people at fully-loaded $500K: $75M–$150M/yr of internal cyber personnel cost across the two leading labs, before compute.
Labs consistently build before they buy at the frontier, so external spend is a fraction of that. At 20–40% of internal — a normal ratio for specialised security work — external cyber evaluation and data spend across the two leading labs plausibly sits at $20M–$60M/yr, and the whole frontier-lab market including DeepMind, Meta, Microsoft and Amazon at $40M–$120M/yr.
The frontier-lab market for externally-sourced cyber evaluation content and expert cyber data in 2026 is very likely between $25M and $120M annually, most likely $40M–$80M. It is concentrated in two to four buyers and currently served by two named incumbents with roughly $120M of combined venture funding and fewer than 100 employees between them.
Two methods built on entirely different inputs — vendor capitalisation and expert-labour cost — landing in the same band is mild comfort, not proof. Both could be wrong in the same direction.
The four figures the record says not to use
Every one of these will appear in a competitor's deck. Using any of them in front of a sophisticated buyer costs credibility that the real numbers would have bought.
1. UK AISI's "£360m" budget. Widely cited, [UNVERIFIED], no primary source reachable, and inconsistent with other UK AI budget reporting. The verifiable UK figures are ARIA's £20m programme with a £2–3m red-team line, and £90m of new cyber-resilience funding announced alongside the GPT-5.5 evaluation (UK AISI). Use those.
2. Project Glasswing's "$100M". From a secondary blog and uncorroborated by Anthropic [UNVERIFIED] (tech-insider.org). What Anthropic actually discloses is partner counts — ~50 initially, ~150 more across 15+ countries — and no money at all (Anthropic).
3. The $2.26bn AI red-teaming services market. Put at $2.26bn in 2026 growing to $6.17bn by 2030 at 28.5% CAGR (Research and Markets) [WEAK] — a paid syndicated report with an uninspectable methodology, whose definition spans "penetration testing, vulnerability assessment, adversarial attack simulation, security auditing." That is mostly conventional security consulting sold to enterprises. It is a different market, and it is gross services billings, not net data revenue (GMV is not revenue). Leading with it invites a discount on everything else you say.
4. The $130–$320/hr OpenAI red-team contractor rates. From an aggregator with no citations. Indicative only, and not a basis for pricing. The rates that are actually posted are Mercor's $70–90/hr for offensive cyber analysis and $85–95/hr for defensive (Mercor) — see Offensive supply and Defensive supply.
What would move the band
Upward. OpenAI's Astra Critical determination happened three weeks before this research and is reflected in none of the numbers above; a lab that halted a training run on cyber grounds and committed publicly to expanding third-party testing will spend materially more over the next twelve months. Anthropic's stated pivot to "harder evaluations" after total benchmark saturation points the same way. The EU AI Act's live adversarial-testing obligation and the June 2026 executive order's classified cyber benchmarking both add compliance-driven volume (Government buyers). And AIxCC's open-sourcing contaminated the free corpus for benchmark purposes, pushing labs toward private held-out content.
Downward. Labs may internalise rather than buy — Anthropic's $485K–$755K cyber lead req is a build signal, not a buy signal. Automated adversarial generation (Gray Swan's Shade, OpenAI's own tooling) may substitute for expert-authored content at the volume tier, leaving only a thin premium tier for humans (Paying the crowd). And a Critical determination that leads to a broad development pause could shrink the training-data market it was supposed to grow.
Scale calibration
The cyber slice is small against the whole. Mercor alone reports $2.0bn annualised gross revenue at roughly a 35% take rate, paying out ~$1.5M/day to 30,000+ vetted experts [WEAK — secondary] (ValueAdd VC). Note that this is gross billings, not net revenue — a distinction What a rake can actually be and GMV is not revenue exist to enforce. One academic estimate puts 2024 data-labelling costs across major companies at roughly 3.1x marginal compute costs for frontier training runs (Kang).
If cyber is 2–5% of frontier expert-data spend — a guess, flagged as one — that is consistent with the $40M–$120M band. Consistent, not confirming.
What a lab pays per hour, or per delivered environment, for expert cyber trajectory data. No public figure exists at any lab. Mercor's posted rates are a cost input, not a sale price, so gross margin on this business cannot be modelled from public information at all. Nothing else in this dossier would improve as much from a single primary conversation.
Also unestablished: whether cyber evaluation spend is growing or flat. No time series exists. The Astra event suggests growth; there is no data to prove it.