Miju Labs

The security dossier

What the labs pay

The dossier's biggest hole, closed three ways: the six firms Anthropic pays are named and one of them publicly sells this exact product to frontier labs; UK government transparency data exposes five real contract values including £702,730 to Gray Swan in a single day; and Mercor's four cyber rates resolve into a clean 3x credential premium that triples our estimate of what a premium environment costs to build.

high confidence10 minupdated 2026-08-30buyers · pricing · contracts · procurement · Mercor · UK AISI · competitors

Until now this dossier could describe demand but not price it. Three findings close that gap, and one of them is bad news.

The six firms are named — and one of them already sells this

Anthropic's coordinated vulnerability disclosure programme has, since May 2026, referred to "one of six independent security research firms" doing the human triage. The Glasswing update of 22 May 2026 does not name them. The names sit one click deeper, on the About page of the CVD dashboard at Anthropic's Frontier Red Team subdomain:

"EXTERNAL SECURITY RESEARCH FIRM PARTNERS — The following are the external security research firms that work with us to triage Claude's vulnerability findings and notify project maintainers. Additional detail on their analysis will follow in subsequent dashboard updates." (red.anthropic.com)

Ada Logics. Anvil. Calif.io. Doyensec. Ophion Security. Trail of Bits.

The same dashboard gives the volume, which is the number that matters commercially. As of 26 August 2026 (overview): 26,153 candidate findings produced by Claude models, 5,008 triaged, and 4,576 reviewed by external security firms at a 91.4% true-positive rate. Of those, 1,022 were confirmed and reported, with a further 1,278 sent direct to maintainers at maintainers' request; 2,300 total disclosed across 392 open-source projects, 1,815 acknowledged, 421 patched, and 462 identifiers issued (177 CVE, 285 GHSA).

Anthropic states plainly that "the process of independent human triage and review is the rate limiting step." Six firms manually reproducing and severity-rating roughly 4,576 candidate vulnerabilities in about six months is a paid, recurring, capacity-constrained human bottleneck with a published supplier list. That is the demand shape The one shape that survives argues for, confirmed by the buyer.

Two of the six are independently observable in the wild. wolfSSL's advisory list carries eight separate advisories credited "Thanks to Calif.io in collaboration with Claude and Anthropic Research for the report", plus three credited to "Max at Trail of Bits" (wolfSSL). Anyone wanting per-firm share of wallet can scrape the credit lines across the 392 listed projects.

Calif.io is the finding that should change the plan

Of the six, one is not a pentest shop absorbing overflow. Calif.io, founded by Thai Duong of BEAST and CRIME fame, states on its own site: "We partner with frontier labs and stay with our customers for the long haul. Half our work is securing AI." Its named service lines include:

"Capability evaluation — We benchmark what a model can actually do offensively, measured against real targets rather than capture-the-flag toys, so the claims you publish are ones you can defend."

That is this product. Written by someone else, on a live commercial website, with a customer wall carrying Google DeepMind, Google Cloud, Let's Encrypt, Figure AI, Cresta AI, Applied Compute, Recall AI, Capital Group and Jump Trading, and testimonials from Jason Clinton (Deputy CISO, Anthropic) and Joel Weinberger (Security, Anthropic), plus Travis McPeak at Cursor, Jim Higgins at CoreWeave and Andy Nguyen at Google (calif.io).

A direct competitor is already selling capability evaluation to at least two frontier labs. Not planning to; selling. That is the single most important competitive fact in this dossier and it did not appear in the earlier competitive map alongside Gray Swan, in full and Irregular, in full.

The other five are more conventional. Trail of Bits is the incumbent US firm and also a UK AISI supplier. Ada Logics (UK/Denmark, founded by researchers from Oxford and Copenhagen) is a long-standing OSS-Fuzz and OpenSSF audit contractor, which explains its fit for open-source triage at volume. Doyensec is an offensive firm with a published large-language-models service line. Anvil Secure is an employee-owned US pentest firm with a named AI Security Testing practice. Ophion Security is the outlier — primarily a product company selling an automated attack-surface platform, and the only one of the six with a published pricing page, which makes it the closest thing to a public price for AI-assisted vulnerability triage anywhere.

Contract values now exist

The earlier dossier opened its buyer analysis with "No contract values, anywhere. Not one lab-to-evaluator contract value is public." Retire that claim.

The source is not a contracts portal. It is DSIT's monthly transparency publication of spend over £25,000. UK AISI sits inside DSIT, so its supplier payments appear line by line with cost-centre labels — including a dedicated "Safety – Cyber" code (collection index).

SupplierTotal foundDetail
Gray Swan Security Inc£702,730Two invoices, both 31 March 2025 — £64,385 and £638,345 — under Safety – Cyber – Consultancy
Irregular (as Pattern Labs Tech Inc)£459,000£260,100 (30 Mar 2025), £153,000 (4 Apr 2025), £45,900 (31 Jul 2025)
Trail Of Bits Inc£245,967£114,298.75 (Jan 2024), £52,059.92 (Apr 2024), £79,608.75 (May 2024)
Crystal Peak Security LLC£116,66615 March 2024
Hack the Box Ltd£39,3669 July 2025, Safety – Cyber – Consultancy

Three things follow.

Gray Swan billed one government buyer £702,730 in a single day. That is materially larger than anything the inference-based sizing in Sizing the cyber pot assumed for a single non-lab customer — and it is a government, not a frontier lab. Against ~50 staff and expansion-led six-figure deals, it is a substantial share of a small company's revenue. See Gray Swan, in full.

Irregular's £459,000 across three payments in 2025 gives the first hard anchor on what a single public-sector customer is worth to a company with $80M raised and roughly 35 staff. See Irregular, in full.

Hack The Box's £39,366 is the best per-environment cross-check available. It corresponds to the "Cooling Tower" ICS range, which took roughly 15 human-hours to solve. At the 10–40x build-to-solve ratio used in The oracle decides everything, that implies 150–600 build hours and a blended £66–260/hour — which brackets the elite contractor tier below almost exactly.

SpecterOps is absent, and the reason is timing, not secrecy. Its 32-step "The Last Ones" corporate range is cited in AISI's 2026 model evaluations, but DSIT's transparency publication currently stops at September 2025; the 2025 file set was published in July 2026 and no 2026 file exists. A grep across all 21 monthly files returns zero SpecterOps matches. An FOI to DSIT would close it — payments over £25k are already routinely published, so the exemption argument is weak. One caveat for anyone modelling from this data: entries labelled "Supplier name withheld" across the AISI rows total roughly £1.79m, so some vendor payments are permanently invisible here. Government buyers has the procurement shape.

Separately, and from the seller side: Bugcrowd's Reinforcement Learning Engineer (Cybersecurity) posting names its buyers outright. The RL and Reasoning Team "design[s] pipelines that ingest software projects, analyze them with Bugcrowd's Mayhem platform, and automatically construct training environments used by frontier AI labs including Anthropic, OpenAI, and Cohere" (Greenhouse, $176,400–$242,550). That is a named, currently-transacting customer list for the adjacent product — and it adds Cohere to the buyer map, which appears nowhere else in this dossier. See The bounty platforms.

The Mercor rate conflict resolves into a 3x credential premium

Four different Mercor cyber rates have been circulating and they looked contradictory. Read out of Mercor's embedded JSON payload — which exposes structured rate, status and creation fields rather than rendered marketing copy — all four are real, current, and describe four different jobs.

RateListingWorkGating
$70–90/hrCybersecurity Experts (31 Mar 2026, private)Vulnerability classification and pattern recognition for "a leading AI lab"; low-level C/C++/Java2+ years only — and the only listing carrying US/CA/UK/AU/NZ residence plus an enhanced background check
$80–90/hrCybersecurity Expert (5 Aug 2026, open)GRC and SOC evaluation scenarios and rubrics: NIST CSF and SOC 2 on a US track, ISO 27001 and NIS2 on an international trackCISSP/CISM preferred; 5+ years as security engineer or CISO
$85–95/hrCyber Security Experts (27 Feb 2026, filled)Blue-team incident analysis plus red-team attack analysis for "a cutting-edge AI research lab"Senior practitioner, both sides
$200–250/hrCybersecurity Research Expert — Offensive Security & Vulnerability Research (4 Aug 2026)"Evaluate AI-generated analyses of complex cybersecurity scenarios, exploits, and vulnerability reports… Contribute to benchmark development"Multiple of: top-ranked CTF team member, CVE discoverer, publicly recognised bug bounty researcher. Part-time, no geographic restriction

The resolution, stated plainly: the market pays a 2.5–3x premium for demonstrated elite offensive credentials over generalist senior security practice. $70–90 buys a competent engineer classifying vulnerabilities. $80–95 buys a senior SOC or GRC practitioner authoring scenarios and rubrics — see Governance, risk and compliance. $200–250 buys someone who has actually shipped CVEs or placed in top-tier CTFs, and the task at that tier is explicitly benchmark construction and expert grading of model-generated exploit reasoning — see Vulnerability research.

Note also what the gating tracks. The residence requirement and enhanced background check land on the cheapest listing, not the most expensive. Vetting scales with the sensitivity of the corpus the contributor touches, not with seniority — which corrects an earlier reading in Copy Daybreak's architecture that no background checks were imposed.

The top tier pays more than the vendors pay their own staff

Gray Swan's Red Team Engineer band is $110,000–$185,000. At the top of band, fully utilised, that is about $89/hour before overhead — almost exactly Mercor's senior contractor rate and less than half the elite contractor rate. Bugcrowd's Cleared Vulnerability Research Engineer runs $154,800–$193,500; HackerOne's Staff Software Engineer, Applied AI runs $230,000–$280,000.

Contract elite offensive talent is priced above what the specialist vendors pay their own permanent staff. That is the clearest single signal in this dossier that scarce credentialed offensive capability — not compute, not tooling, not demand — is the binding constraint in this market. Offensive supply and Paying the crowd work through what that means for recruiting.

The correction that follows

Earlier sizing in this dossier used $70–90/hr to compute "$15K–$70K of raw expert labour per premium environment". If a premium environment is built by the credentialed tier — and by definition a premium one is — that figure is wrong by roughly a factor of three.

Revised estimate

$40K–$200K of raw expert labour per premium environment, replacing the earlier $15K–$70K. The multiplier is the 2.5–3x credential premium; the independent cross-check is Hack The Box's £39,366 range at an implied £66–260/hour. Treat the earlier figure as retired.

What is still missing

Per-hour is the only unit anyone quotes. No lab, vendor or marketplace publishes a per-trajectory, per-task or per-environment price. The nearest per-environment figures in existence are the Hack The Box invoice above and DARPA AIxCC's "~$152 per competition task" for a fully automated system — a floor, not a comparator.

No lab publishes a vendor security standard. OpenAI's trust portal lists what OpenAI holds — SOC 2 Type 2, ISO 27001/27017/27018/27701/42001, CSA STAR, FedRAMP 20x, PCI DSS — and specifies nothing required of suppliers. Anthropic's trust centre is behind a JavaScript app. The real constraints arrive through the Mercor listings above and through Anthropic's Cyber Verification Program, whose conditions include that organisations on Zero Data Retention are not eligible — a documented operational problem for any vendor promising its own buyers ZDR.

Per-firm volumes for the six triage partners are not published, though Anthropic says "additional detail on their analysis will follow in subsequent dashboard updates". That update would convert a name list into a share-of-wallet estimate, and it is worth watching for.

Read this page against The labs as buyers for demand, Government buyers for the procurement route, The measurement gap for what the money is buying, and The security read for whether any of it adds up to a company.