Miju Labs

The security dossier

The labs as buyers

Two buyers hold the money, and both have said in writing that their cyber evaluation stock is exhausted — one paused a frontier training run over it, the other published the sentence that is the whole sales pitch.

high confidence9 minupdated 2026-08-30frontier labs · cyber · buyers · safety frameworks

Two things happened in 2026 that reframe every question about who pays for cyber evaluation content.

On 7 August 2026 OpenAI announced that its upcoming model Astra had reached the Critical cybersecurity capability level under its own Preparedness Framework, and paused (OpenAI). Per Axios it halted two weeks of deployment-focused RL training, put its largest planned frontier RL run on hold, left "a significant number of Astra and cyber-related research workloads" paused, and began rewriting the framework itself (Axios, 18 Aug 2026). It promised "recommended security controls to third-party testing partners for running higher risk evaluations."

The second is a sentence Anthropic wrote about itself. The Claude Opus 4.6 system card states that the model "has saturated all of our current cyber evaluations, achieving ~100% on Cybench (pass@30) and 66% on CyberGym (pass@1)," and that "[t]he saturation of our evaluation infrastructure means we can no longer use current benchmarks to track capability progression" (system card). Anthropic's stated response was to prioritise "investment in harder evaluations and enhanced monitoring for cyber misuse, even in the absence of formal RSP thresholds." By Claude Opus 5, CyberGym is listed as retired.

That second sentence is the sales pitch, written by the buyer. See The specialist wedge for why an unclaimed public artefact matters more than a deck.

The frameworks, and where cyber sits in them

LabFrameworkCyber thresholdStatus, Aug 2026
OpenAIPreparedness Framework v2, 15 Apr 2025High and Critical, both defined verbatimAstra assessed Critical; GPT-5.5 "High, below Critical"
AnthropicRSP v3.0None at any ASLCyber evals saturated; buying harder ones anyway
Google DeepMindFrontier Safety Framework v3.1, 17 Apr 2026Cyber Uplift Level 1Gemini 3.7 Flash at alert threshold; Key Skills Benchmark not run
MetaFrontier AI FrameworkCyber 1, 2, 3 with named threat scenariosCyber 1 trigger written as "complex CTF challenges at the level of a professional cyber expert"
Microsoft AIFrontier Governance Framework, Feb 2026Four levels, Low to CriticalCommits to "engage qualified third parties"
xAICould not establishNone foundCorporate infosec only
MistralCould not establishNone foundOut of scope

Sources: OpenAI PF v2, Anthropic RSP v3.0, DeepMind FSF, Gemini 3.7 Flash report, Meta via ETO AGORA, Microsoft FGF.

The counter-intuitive line is Anthropic's. RSP v3.0 covers non-novel CBRN, novel CBRN, high-stakes sabotage and automated R&D. The Opus 4.6 system card says it plainly: "The RSP does not define a formal capability threshold for cyber risks at any AI Safety Level." Anthropic still outspends everyone on cyber — the highest posted band in the industry, a Frontier Red Team cyber sub-team, a Cybersecurity Products org, a Safeguards cyber-harms stack, and a cyber-specialised model line.

Policy documents are not the demand signal. Payroll is.

Every cyber-titled posting found

Pulled from the labs' own applicant-tracking APIs on 30 August 2026 — Anthropic and xAI via Greenhouse, OpenAI via Ashby. Anthropic's board carried 571 open roles that day, OpenAI's 757, xAI's 252. Roles sharing a band are grouped.

TitleLabPosted band (USD/yr)Source
Lead, Frontier Red Team (Cyber)Anthropic$485,000–$755,000gh
RS, Frontier Red Team (Emerging Risks)Anthropic$320,000–$850,000gh
RE/RS, Frontier Red Team (Cyber), closedAnthropic$320,000–$485,000gh
ML/Research Engineer, ML Infra Engineer, SafeguardsAnthropic$350,000–$500,000gh
Cybersecurity Products: researcher, SWE, EMAnthropic$405,000–$485,000gh
Threat Intel Manager, Model ExploitationAnthropic$375,000–$455,000gh
Head of Vulnerability DisclosureAnthropic$330,000–$395,000gh
Staff+ SWE, Safeguards Evals and DataAnthropic$320,000–$485,000gh
Red Team Engineer, Threat Intelligence EngineerAnthropic$320,000–$405,000gh
Research Engineer, Cybersecurity RLAnthropic$300,000–$405,000gh
Product Manager, Cybersecurity; PM, Safeguards (Cyber)Anthropic$305,000–$460,000gh
Cyber Harms: Enforcement Lead, Policy Lead, Enforcement AnalystAnthropic$285,000–$330,000gh
Cyber AE, Applied AI Architecture Mgr, Partner Dev, Policy AnalystAnthropic$190,000–$450,000gh
Cyber Threat Investigator; AI Security FellowsAnthropic$230,000–$290,000; $3,850/wkgh
Data Scientist, CybersecurityOpenAI$263,000–$515,000ashby
Offensive Security Engineer, Agent Engineer, Principal SWE Codex CyberOpenAI$347,000–$490,000ashby
Head of Government Cyber IntegrationOpenAI$401,000–$445,000ashby
Researcher, Frontier Cybersecurity Risks (Safety Systems)OpenAI$295,000–$445,000ashby
Agentic AI Threats researcher; Preparedness Lead, Coding AgentsOpenAI$293,000–$405,000ashby
Data Scientist, PreparednessOpenAI$347,000–$400,000ashby
SWE and Full Stack SWE, Codex CyberOpenAI$230,000–$405,000ashby
Blue-team PM, Codex controls PM, threat investigator, agent securityOpenAI$230,000–$385,000ashby
TPM, audit, partnerships, applied AI, sales, comms and PMM — seven cyber rolesOpenAI$221,000–$445,000ashby
Model Policy, Frontier Cyber RiskOpenAI$266,000–$335,000ashby
Red Team Specialist — Cyber (dataset construction)OpenAI$198,000–$320,000ashby
Cyber Operations Strategist, Critical Harm Ops (Ontario)OpenAICA$140,000–CA$188,000ashby
Research Engineer, Frontier Safety Risk AssessmentGoogle DeepMind$136,000–$245,000gh
Three further Frontier Safety / AGI Safety rolesGoogle DeepMindnot postedGoogle Careers
Security Engineer, AI Security Frontier RisksMeta$177,000–$251,000mirror
Four corporate infosec and GRC rolesxAI$100,000–$258,000gh
Expression of Interest — Red TeamUK AISI£65,000–£145,000 + 28.97% pensiongh

The money is in two accounts. Anthropic and OpenAI carry roughly 80 cyber- and security-titled open roles between them. DeepMind's frontier-safety function is four roles; Meta's is one. Anthropic's cyber lead band is roughly three times DeepMind's posted frontier-safety band and 2.7x Meta's. See One customer is a binary event.

The reqs describe the product. The Anthropic cyber lead posting lists "[t]raining models to autonomously find and patch vulnerabilities using tens of trillions of tokens" and "[p]ointing autonomous AI systems at real-world security challenges (bug bounties, CTFs etc.) to characterize risks… and compare to human experts" (posting). Tens of trillions of cyber RL tokens plus human expert baselines is a purchase order written as a job description.

xAI is not a buyer. Its entire security posture is corporate infosec and GRC plus two unbanded threat-intelligence analysts. It runs a large Human Data organisation and hires AI tutors in ~30 languages — with no cyber vertical inside it and no frontier cyber-evaluation function. Do not build a pipeline for xAI.

Who is already credited in the system cards

Model cardLabExternal party credited for cyber
Claude Opus 5, Jul 2026AnthropicIrregular (CyScenarioBench), UK AISI, Mozilla, Trajectory Labs PBC (100 hrs), 10a Labs (16 hrs), Gray Swan
Claude Opus 4.8, May 2026AnthropicNone named in cyber; Mozilla collaboration, OSS-Fuzz corpus
Claude Opus 4.6, Feb 2026AnthropicUS CAISI
Claude Mythos PreviewAnthropicProfessional human security contractors (198 reports manually validated)
GPT-5 / o3 / o4-miniOpenAIIrregular
GPT-5.5OpenAIIrregular, US CAISI, UK AISI
GPT-5.6-CyberOpenAISpecterOps cited as user; advisers CrowdStrike, METR, Redwood Research
Gemini 3 Pro, Nov 2025Google DeepMindExternal testers referenced, not named
Gemini 3.7 Flash, Aug 2026Google DeepMindGray Swan (Shade agent, 28 cyber tasks)
Meta system cardsMetaGray Swan

Sources: Opus 5 system card, Mythos Preview, OpenAI safety hub, Gemini 3.7 Flash report, Gray Swan release. Earlier Claude 4 and 3.7 Sonnet cards credit Pattern Labs, now Irregular.

Two names carry that table. Irregular runs on roughly 35 people and appears in cyber sections at four labs — "one vendor whose methodology, benchmarks and infrastructure are load-bearing for multiple competitors' safety claims at once" (BERI). Gray Swan AI appears in eleven system cards and sells breadth, not depth. Neither sells what the Mythos write-up says is missing: humans who can validate logic vulnerabilities, where Anthropic states it "lose[s] the ability to (near-)perfectly validate" without a crash oracle. That is The one shape that survives.

The gated cyber model lines

Anthropic's Claude Mythos is "[o]ur most capable model for cybersecurity and biology research," at $10/M input and $50/M output, open only to a small set of testing partners under mandatory 30-day retention (Anthropic). Its distribution arm Project Glasswing started at ~50 critical-infrastructure and open-source partners and added ~150 more across 15+ countries (Anthropic). OpenAI's GPT-5.6-Cyber is a reduced-refusal fine-tune at $12.50/M input and $75/M output, gated behind "Daybreak Red" — demonstrated authorised defensive work plus SOC 2 Type II or ISO 27001, SSO, MFA and monitoring (VentureBeat).

So what

A dedicated cyber model line implies a dedicated cyber training-data line. Two labs now ship separately-priced, separately-gated cyber models. Those were post-trained on something, and it was not a general web corpus.

The procurement process nobody has sized

Microsoft commits in writing to "engage qualified third parties to conduct evaluations in ways that are appropriate to the risk profile," with selection made by "internal experts who are independent from the model development team" (FGF). That is a vendor-selection mechanism written into a safety framework — and Microsoft has no visible cyber-evaluation hiring, which reads as intent to buy rather than build.

What could not be established

No lab-to-evaluator contract value is public anywhere — not Irregular's, not Gray Swan's, not Trajectory Labs' or 10a Labs'. OpenAI confirms it pays external testers via "direct payment and/or API credits" and that "[n]o payment is ever contingent on the results" (OpenAI), without amounts. Microsoft AI's and Mistral's job boards could not be enumerated. Project Glasswing's size is undisclosed; the "$100M" figure on a secondary blog is [UNVERIFIED]. See Sizing the cyber pot.

The Frontier Model Forum — covering thresholds from Amazon, Anthropic, Google, Meta, Microsoft and OpenAI — concedes that third-party "evaluation maturity remains uneven" and that the ecosystem needs "standardized cyber evaluation frameworks, shared red-teaming methodologies, and common benchmarks" (FMF). Six labs jointly calling their own market immature is an entry invitation in writing.

Next: Government buyers, Commercial buyers where the evidence runs out, and Offensive supply for who you would hire.