Two things happened in 2026 that reframe every question about who pays for cyber evaluation content.
On 7 August 2026 OpenAI announced that its upcoming model Astra had reached the Critical cybersecurity capability level under its own Preparedness Framework, and paused (OpenAI). Per Axios it halted two weeks of deployment-focused RL training, put its largest planned frontier RL run on hold, left "a significant number of Astra and cyber-related research workloads" paused, and began rewriting the framework itself (Axios, 18 Aug 2026). It promised "recommended security controls to third-party testing partners for running higher risk evaluations."
The second is a sentence Anthropic wrote about itself. The Claude Opus 4.6 system card states that the model "has saturated all of our current cyber evaluations, achieving ~100% on Cybench (pass@30) and 66% on CyberGym (pass@1)," and that "[t]he saturation of our evaluation infrastructure means we can no longer use current benchmarks to track capability progression" (system card). Anthropic's stated response was to prioritise "investment in harder evaluations and enhanced monitoring for cyber misuse, even in the absence of formal RSP thresholds." By Claude Opus 5, CyberGym is listed as retired.
That second sentence is the sales pitch, written by the buyer. See The specialist wedge for why an unclaimed public artefact matters more than a deck.
The frameworks, and where cyber sits in them
| Lab | Framework | Cyber threshold | Status, Aug 2026 |
|---|---|---|---|
| OpenAI | Preparedness Framework v2, 15 Apr 2025 | High and Critical, both defined verbatim | Astra assessed Critical; GPT-5.5 "High, below Critical" |
| Anthropic | RSP v3.0 | None at any ASL | Cyber evals saturated; buying harder ones anyway |
| Google DeepMind | Frontier Safety Framework v3.1, 17 Apr 2026 | Cyber Uplift Level 1 | Gemini 3.7 Flash at alert threshold; Key Skills Benchmark not run |
| Meta | Frontier AI Framework | Cyber 1, 2, 3 with named threat scenarios | Cyber 1 trigger written as "complex CTF challenges at the level of a professional cyber expert" |
| Microsoft AI | Frontier Governance Framework, Feb 2026 | Four levels, Low to Critical | Commits to "engage qualified third parties" |
| xAI | Could not establish | None found | Corporate infosec only |
| Mistral | Could not establish | None found | Out of scope |
Sources: OpenAI PF v2, Anthropic RSP v3.0, DeepMind FSF, Gemini 3.7 Flash report, Meta via ETO AGORA, Microsoft FGF.
The counter-intuitive line is Anthropic's. RSP v3.0 covers non-novel CBRN, novel CBRN, high-stakes sabotage and automated R&D. The Opus 4.6 system card says it plainly: "The RSP does not define a formal capability threshold for cyber risks at any AI Safety Level." Anthropic still outspends everyone on cyber — the highest posted band in the industry, a Frontier Red Team cyber sub-team, a Cybersecurity Products org, a Safeguards cyber-harms stack, and a cyber-specialised model line.
Policy documents are not the demand signal. Payroll is.
Every cyber-titled posting found
Pulled from the labs' own applicant-tracking APIs on 30 August 2026 — Anthropic and xAI via Greenhouse, OpenAI via Ashby. Anthropic's board carried 571 open roles that day, OpenAI's 757, xAI's 252. Roles sharing a band are grouped.
| Title | Lab | Posted band (USD/yr) | Source |
|---|---|---|---|
| Lead, Frontier Red Team (Cyber) | Anthropic | $485,000–$755,000 | gh |
| RS, Frontier Red Team (Emerging Risks) | Anthropic | $320,000–$850,000 | gh |
| RE/RS, Frontier Red Team (Cyber), closed | Anthropic | $320,000–$485,000 | gh |
| ML/Research Engineer, ML Infra Engineer, Safeguards | Anthropic | $350,000–$500,000 | gh |
| Cybersecurity Products: researcher, SWE, EM | Anthropic | $405,000–$485,000 | gh |
| Threat Intel Manager, Model Exploitation | Anthropic | $375,000–$455,000 | gh |
| Head of Vulnerability Disclosure | Anthropic | $330,000–$395,000 | gh |
| Staff+ SWE, Safeguards Evals and Data | Anthropic | $320,000–$485,000 | gh |
| Red Team Engineer, Threat Intelligence Engineer | Anthropic | $320,000–$405,000 | gh |
| Research Engineer, Cybersecurity RL | Anthropic | $300,000–$405,000 | gh |
| Product Manager, Cybersecurity; PM, Safeguards (Cyber) | Anthropic | $305,000–$460,000 | gh |
| Cyber Harms: Enforcement Lead, Policy Lead, Enforcement Analyst | Anthropic | $285,000–$330,000 | gh |
| Cyber AE, Applied AI Architecture Mgr, Partner Dev, Policy Analyst | Anthropic | $190,000–$450,000 | gh |
| Cyber Threat Investigator; AI Security Fellows | Anthropic | $230,000–$290,000; $3,850/wk | gh |
| Data Scientist, Cybersecurity | OpenAI | $263,000–$515,000 | ashby |
| Offensive Security Engineer, Agent Engineer, Principal SWE Codex Cyber | OpenAI | $347,000–$490,000 | ashby |
| Head of Government Cyber Integration | OpenAI | $401,000–$445,000 | ashby |
| Researcher, Frontier Cybersecurity Risks (Safety Systems) | OpenAI | $295,000–$445,000 | ashby |
| Agentic AI Threats researcher; Preparedness Lead, Coding Agents | OpenAI | $293,000–$405,000 | ashby |
| Data Scientist, Preparedness | OpenAI | $347,000–$400,000 | ashby |
| SWE and Full Stack SWE, Codex Cyber | OpenAI | $230,000–$405,000 | ashby |
| Blue-team PM, Codex controls PM, threat investigator, agent security | OpenAI | $230,000–$385,000 | ashby |
| TPM, audit, partnerships, applied AI, sales, comms and PMM — seven cyber roles | OpenAI | $221,000–$445,000 | ashby |
| Model Policy, Frontier Cyber Risk | OpenAI | $266,000–$335,000 | ashby |
| Red Team Specialist — Cyber (dataset construction) | OpenAI | $198,000–$320,000 | ashby |
| Cyber Operations Strategist, Critical Harm Ops (Ontario) | OpenAI | CA$140,000–CA$188,000 | ashby |
| Research Engineer, Frontier Safety Risk Assessment | Google DeepMind | $136,000–$245,000 | gh |
| Three further Frontier Safety / AGI Safety roles | Google DeepMind | not posted | Google Careers |
| Security Engineer, AI Security Frontier Risks | Meta | $177,000–$251,000 | mirror |
| Four corporate infosec and GRC roles | xAI | $100,000–$258,000 | gh |
| Expression of Interest — Red Team | UK AISI | £65,000–£145,000 + 28.97% pension | gh |
The money is in two accounts. Anthropic and OpenAI carry roughly 80 cyber- and security-titled open roles between them. DeepMind's frontier-safety function is four roles; Meta's is one. Anthropic's cyber lead band is roughly three times DeepMind's posted frontier-safety band and 2.7x Meta's. See One customer is a binary event.
The reqs describe the product. The Anthropic cyber lead posting lists "[t]raining models to autonomously find and patch vulnerabilities using tens of trillions of tokens" and "[p]ointing autonomous AI systems at real-world security challenges (bug bounties, CTFs etc.) to characterize risks… and compare to human experts" (posting). Tens of trillions of cyber RL tokens plus human expert baselines is a purchase order written as a job description.
xAI is not a buyer. Its entire security posture is corporate infosec and GRC plus two unbanded threat-intelligence analysts. It runs a large Human Data organisation and hires AI tutors in ~30 languages — with no cyber vertical inside it and no frontier cyber-evaluation function. Do not build a pipeline for xAI.
Who is already credited in the system cards
| Model card | Lab | External party credited for cyber |
|---|---|---|
| Claude Opus 5, Jul 2026 | Anthropic | Irregular (CyScenarioBench), UK AISI, Mozilla, Trajectory Labs PBC ( |
| Claude Opus 4.8, May 2026 | Anthropic | None named in cyber; Mozilla collaboration, OSS-Fuzz corpus |
| Claude Opus 4.6, Feb 2026 | Anthropic | US CAISI |
| Claude Mythos Preview | Anthropic | Professional human security contractors (198 reports manually validated) |
| GPT-5 / o3 / o4-mini | OpenAI | Irregular |
| GPT-5.5 | OpenAI | Irregular, US CAISI, UK AISI |
| GPT-5.6-Cyber | OpenAI | SpecterOps cited as user; advisers CrowdStrike, METR, Redwood Research |
| Gemini 3 Pro, Nov 2025 | Google DeepMind | External testers referenced, not named |
| Gemini 3.7 Flash, Aug 2026 | Google DeepMind | Gray Swan (Shade agent, 28 cyber tasks) |
| Meta system cards | Meta | Gray Swan |
Sources: Opus 5 system card, Mythos Preview, OpenAI safety hub, Gemini 3.7 Flash report, Gray Swan release. Earlier Claude 4 and 3.7 Sonnet cards credit Pattern Labs, now Irregular.
Two names carry that table. Irregular runs on roughly 35 people and appears in cyber sections at four labs — "one vendor whose methodology, benchmarks and infrastructure are load-bearing for multiple competitors' safety claims at once" (BERI). Gray Swan AI appears in eleven system cards and sells breadth, not depth. Neither sells what the Mythos write-up says is missing: humans who can validate logic vulnerabilities, where Anthropic states it "lose[s] the ability to (near-)perfectly validate" without a crash oracle. That is The one shape that survives.
The gated cyber model lines
Anthropic's Claude Mythos is "[o]ur most capable model for cybersecurity and biology research," at $10/M input and $50/M output, open only to a small set of testing partners under mandatory 30-day retention (Anthropic). Its distribution arm Project Glasswing started at ~50 critical-infrastructure and open-source partners and added ~150 more across 15+ countries (Anthropic). OpenAI's GPT-5.6-Cyber is a reduced-refusal fine-tune at $12.50/M input and $75/M output, gated behind "Daybreak Red" — demonstrated authorised defensive work plus SOC 2 Type II or ISO 27001, SSO, MFA and monitoring (VentureBeat).
A dedicated cyber model line implies a dedicated cyber training-data line. Two labs now ship separately-priced, separately-gated cyber models. Those were post-trained on something, and it was not a general web corpus.
The procurement process nobody has sized
Microsoft commits in writing to "engage qualified third parties to conduct evaluations in ways that are appropriate to the risk profile," with selection made by "internal experts who are independent from the model development team" (FGF). That is a vendor-selection mechanism written into a safety framework — and Microsoft has no visible cyber-evaluation hiring, which reads as intent to buy rather than build.
No lab-to-evaluator contract value is public anywhere — not Irregular's, not Gray Swan's, not Trajectory Labs' or 10a Labs'. OpenAI confirms it pays external testers via "direct payment and/or API credits" and that "[n]o payment is ever contingent on the results" (OpenAI), without amounts. Microsoft AI's and Mistral's job boards could not be enumerated. Project Glasswing's size is undisclosed; the "$100M" figure on a secondary blog is [UNVERIFIED]. See Sizing the cyber pot.
The Frontier Model Forum — covering thresholds from Amazon, Anthropic, Google, Meta, Microsoft and OpenAI — concedes that third-party "evaluation maturity remains uneven" and that the ecosystem needs "standardized cyber evaluation frameworks, shared red-teaming methodologies, and common benchmarks" (FMF). Six labs jointly calling their own market immature is an entry invitation in writing.
Next: Government buyers, Commercial buyers where the evidence runs out, and Offensive supply for who you would hire.