The most useful thing DARPA did for this business was give away $9.9m of cyber reasoning systems for free.
The AI Cyber Challenge concluded at DEF CON 33 in August 2025. Team Atlanta took $4M, Trail of Bits' Buttercup $3M, Theori $1.5M, from a final-round pool of $8.5M, with ARPA-H adding $1.4M for real-world integration. The systems found 54 of 63 synthetic vulnerabilities (86%, up from 37% at semifinals) and patched 43 of 54, at an average of ~$152 per competition task. Anthropic, Google and OpenAI each donated $350,000 in LLM credits. And all seven finalist systems were open-sourced under OSI-approved licences, along with the competition framework, the challenges, the telemetry and the tooling (DARPA).
The tools kept working afterwards. FuzzingBrain found 62 vulnerabilities across 26 projects; 42-b3yond-6ug found 12 Linux-kernel vulnerabilities plus 10 userspace zero-days; Team Atlanta's OSS-CRS found 25 across 16 projects. OSS-CRS and FuzzingBrain now live at OpenSSF inside the Linux Foundation, which has stood up a Cyber Reasoning Systems Special Interest Group (OpenSSF).
That cuts two ways, and both directions matter.
Anyone selling vulnerability-discovery tooling now competes with free, permanently. Anyone selling novel, unpublished, expert-authored evaluation content just got more valuable, because the open corpus is now public — and therefore contaminated for benchmarking a model that may have trained on it. Open-sourcing AIxCC raised the price of private, held-out, expert-generated evaluation data. It is the strongest structural argument in this whole dossier.
Could not establish a named DARPA successor programme. DARPA says it is "working with public/private sector partners on widespread technology adoption," which is transition language, not a solicitation.
UK AISI is the clearest government buyer, and it subcontracts by name
This is the single best-evidenced government purchase of external cyber evaluation content anywhere in the record, and it names its suppliers.
For its GPT-5.5 cyber evaluation, UK AISI ran 95 narrow cyber tasks across four difficulty tiers plus two cyber ranges — "The Last Ones," a 32-step corporate network simulation estimated at ~20 human hours to solve, and "Cooling Tower," a 7-step ICS simulation at ~15 human hours. It did not build them itself. SpecterOps built The Last Ones. Hack The Box built the Cooling Tower ICS range. Crystal Peak Security and Irregular worked on advanced task design. Expert red-teaming of GPT-5.5's safeguards found a universal jailbreak in six hours (UK AISI).
Four named subcontractors on one evaluation. That is a procurement pattern a new supplier can actually apply to, unlike the lab relationships in The labs as buyers, which are invisible until a system card is published.
The reason the buying continues is in AISI's own trend work: apprentice-level cyber task success moved from under 9% to ~50% between late 2023 and 2025, the first expert-level task completion happened in 2025, and unassisted task duration is "doubling roughly every eight months" (Frontier AI Trends Report). AISI has evaluated more than 30 frontier systems in two years. Its own Red Team hires at £65,000–£145,000 plus a 28.97% pension across Alignment, Misuse and Control, and shares findings with Anthropic, OpenAI and DeepMind (listing). A separate £90m of new UK funding to boost cyber resilience was announced alongside the GPT-5.5 publication.
ARIA: the only precise public unit price in the sector
The UK's Advanced Research and Invention Agency opened Safeguarded AI: Cybersecurity on 20 May 2026 at £20 million total, split into two tracks. Track 1 funds three to six blue teams at roughly £2.5–3.5m each. Track 2 funds one adversarial evaluation partner — a red team — at approximately £2–3 million (ARIA solicitation).
The shape of the work is specified in unusual detail. The red team runs 8-week sprint cycles — six weeks build and verify, one week attack, one week review — producing reports with "reproducible artifacts for any successful breaks." The programme also builds a stakeholder network of cybersecurity practitioners and critical-infrastructure operators to "ground specifications and threat models in real operational conditions." Submissions closed 1 July 2026, notifications went out 31 July 2026, kickoff 1 September 2026, concluding November 2027.
£2–3m for roughly 14 months of one red team. That is the only precise, public unit price for government-procured AI adversarial cyber evaluation in the record. Everything in Sizing the cyber pot is calibrated against it, and any pricing conversation should start there rather than at a per-hour rate.
The award was decided by 31 July 2026 and the winner could not be established. If it went to an existing frontier-lab evaluation vendor, the concentration problem described in One customer is a binary event extends into UK public money. If it went to a security consultancy, the market is broader than the system cards suggest. This is a cheap fact to obtain and it changes the read.
CAISI, honestly
The US Center for AI Standards and Innovation — the renamed AI Safety Institute inside NIST — describes itself as "industry's primary point of contact within the U.S. government to facilitate testing and collaborative research," and leads "evaluations and assessments of capabilities of U.S. and adversary AI systems" (NIST). Its published cyber work includes a joint UK AISI/CAISI ExploitBench assessment of Kimi K3 (July 2026), a DeepSeek V4 Pro evaluation (May 2026), a GLM-5.2 assessment, and the earlier DeepSeek models evaluation. On 6 May 2026 it signed pre-deployment review agreements with Google DeepMind, Microsoft and xAI covering cyber, biosecurity and chemical-weapons risks; OpenAI separately gave early access to GPT-5.5.
Now the honest part. CAISI's named external-engagement mechanisms are federal hiring, guest researcher arrangements, nonprofit collaborations and a CRADA with OpenMined signed March 2026 (ExecutiveGov). CRADAs and guest-researcher arrangements are typically no-cost or cost-shared. No CAISI contract vehicle, RFP or awarded contract for external cyber red teaming could be established; sam.gov search results were not reachable to confirm either way.
Treat CAISI as a demand amplifier, not a buyer. It forces labs to hold evaluations, which grows the labs' spend. It does not itself appear to write cheques for evaluation content. A claim in the other direction would need a contract number attached.
The June 2026 Executive Order on Promoting Advanced Artificial Intelligence Innovation and Security sharpens the same point. Section 3(a) mandates "a classified benchmarking process to assess the advanced cyber capabilities of AI models" used to designate a "covered frontier model," with the Director of NSA making final determinations alongside the National Cyber Director, CISA and DoD (White House). The framework remains voluntary — no licensing, no mandatory preclearance — but developers face pressure to grant 30 days of pre-release access, and the "covered frontier model" definition is itself classified, creating uncertainty about which systems trigger federal engagement (Holland & Knight).
A classified benchmark is bad news and good news. Bad: you cannot sell into it without clearances and a facility. Good: labs preparing for an opaque classified assessment will over-invest in unclassified private evaluation coverage rather than be surprised. Could not establish any funded procurement flowing from the EO as of 30 August 2026.
The EU AI Act: a compliance buyer, with a soft ceiling
Article 55 requires providers of general-purpose AI models with systemic risk to "perform model evaluation in accordance with standardised protocols and tools reflecting the state of the art, including conducting and documenting adversarial testing of the model with a view to identifying and mitigating systemic risks," to assess and mitigate systemic risks at Union level, to report serious incidents to the AI Office "without undue delay," and to "ensure an adequate level of cybersecurity protection" for the model and its physical infrastructure (EU AI Act).
The timing matters more than the text. GPAI obligations became applicable 2 August 2025 and are already enforceable; the systemic-risk trigger is 10^25 FLOP of cumulative training compute or AI Office designation. The Digital Omnibus on AI, given final Council approval on 29 June 2026, did not defer GPAI obligations — it deferred only Annex III high-risk systems to December 2027 and Annex I product-embedded systems to August 2028 (Legiscope).
Two caveats a seller must state before a buyer does. Article 55 requires adversarial testing of systemic risks generally — it does not name cyber as a mandatory test domain, so cyber-specific spend is inferred, not mandated. And until harmonised standards exist, "standardised protocols and tools reflecting the state of the art" is undefined, which hands the compliance buyer wide discretion about how much to spend.
The EU AI Act creates a floor of obligation and a ceiling of ambiguity. It is a genuine demand driver and a soft one. Nobody has been fined for insufficient cyber adversarial testing, and until somebody is, this is a reason a procurement committee says yes rather than a reason it starts.
What a seller should take from this section
Government here is three different things wearing one label. UK AISI and ARIA are real buyers with named suppliers and a published unit price — the only part of this market where you can see the shape of a contract before you win one. CAISI and the June 2026 EO are amplifiers, growing the labs' spend without spending much themselves. DARPA is a competitor that has already shipped, having converted public money into free tooling and a contaminated public corpus — which raises the value of what is not public.
Compare with Commercial buyers, where the evidence for a paying segment largely does not exist, and The labs as buyers, where the money is but the contracts are invisible.