Miju Labs

The security dossier

What the dossier could not establish

A follow-up pass closed the top item — the six firms are Ada Logics, Anvil, Calif.io, Doyensec, Ophion Security and Trail of Bits — and produced the first contract values this market has ever had in public. What is left is re-ranked around a single unknown: what a lab pays, per anything.

high confidence9 minupdated 2026-08-30gaps · open questions · contradictions · research agenda

Everything else in this dossier is an argument built on evidence of varying quality. This is where the evidence runs out — updated after a follow-up pass that closed four of the seven items below and half-closed two more. Answered items are kept, with the answer and the source, because the shape of an answer is worth as much as the fact. The general version is What we could not establish; nothing below duplicates it.


Answered

✔ Who are the six external security research firms?

ANSWERED. They are Ada Logics, Anvil, Calif.io, Doyensec, Ophion Security and Trail of Bits.

The list is on the "About this dashboard" page of Anthropic's coordinated-disclosure site, one click below the announcements that describe only "one of six independent security research firms": "The following are the external security research firms that work with us to triage Claude's vulnerability findings and notify project maintainers" (Anthropic CVD — about). Two are independently corroborable through advisory credits — wolfSSL publishes eight advisories thanking "Calif.io in collaboration with Claude and Anthropic Research" and three thanking "Max at Trail of Bits" (wolfSSL).

What the answer changed. It was supposed to resolve into a market-entry plan, an acquisition list or a hiring list. It resolved into a fourth thing nobody predicted: a direct product competitor. Calif.io — founded by Thai Duong — publishes "Capability evaluation… we benchmark what a model can actually do offensively, measured against real targets rather than capture-the-flag toys" as a service line, says "half our work is securing AI", carries Google DeepMind on its logo wall and quotes Anthropic's Deputy CISO in its testimonials (calif.io). That is the position at The one shape that survives, occupied. It is why room drops to 2 at The security read.

Still open underneath it: per-firm volumes. Anthropic says "additional detail on their analysis will follow in subsequent dashboard updates," which would convert a name list into a share-of-wallet estimate. The dashboard also exposes machine-readable payload.json and ledger.json files whose by_project field can be cross-referenced against each project's advisory credits to attribute findings per firm.

✔ The UK AISI subcontract values

ANSWERED for two of three, plus three bonus values, and the source was not where anyone was looking. Not Contracts Finder, not Find a Tender — DSIT's monthly transparency publication of spend over £25,000, which carries a dedicated "Safety – Cyber" cost centre (DSIT).

Gray Swan Security Inc £702,730 (two invoices, 31 March 2025) · Irregular / Pattern Labs Tech Inc £459,000 (three payments, 2025) · Trail of Bits Inc £245,967 (three payments, 2024) · Crystal Peak Security LLC £116,666 (March 2024) · Hack The Box Ltd £39,366 (July 2025).

What the answer changed. It retires the flat claim that no contract value existed anywhere in this market. It confirms government as a real second buyer at meaningful size. And the Hack The Box line gives the only public per-environment number in existence: £39,366 for the "Cooling Tower" ICS range, which takes ~15 human-hours to solve, implying £66–260/hour blended at the 10–40x build-to-solve ratio in The oracle decides everything.

Still open: SpecterOps. It appears in none of the twenty-one DSIT files from January 2024 to September 2025, because the series ends at September 2025 — the "The Last Ones" range is cited in a 2026 AISI blog. Note too that rows marked "Supplier name withheld" total roughly £1.79m, so some cyber payments may be permanently invisible here. An FOI to DSIT for FY2025-26 and FY2026-27 payments to SpecterOps, Crystal Peak and Hack The Box is the route; payments over £25k are already routinely published, so the exemption argument is weak.

✔ Reset time and cost for a multi-host range

MECHANISM ANSWERED; the conclusion inverts the question. AISI's cyber ranges are not bare metal. They are Proxmox VE virtual machines orchestrated by an open-source Inspect plugininspect-proxmox-sandbox, publicly readable — which "allows Inspect to create and manage a fleet of virtual machines" (AISI incident report). Reset is a clone from a baked template, live QEMU snapshotting including running processes is supported, VMs boot concurrently by default, and throughput is bounded by pool size rather than reset time.

The cost falls out of AISI's own EC2 provisioning path: an m8i.2xlarge with nested virtualisation at roughly $0.423/hr, plus a 1,024 GB gp3 volume at ~$0.11/hr, is about $0.53/hour per concurrency slot — call it $2–8/hour for a realistic multi-host range. A full campaign of 122 runs across two ranges and seven models over four days is a few thousand dollars of infrastructure.

What the answer changed. The question assumed infrastructure might be the constraint. It is not. Infrastructure is negligible; expert build time is the whole cost, which is precisely why the $200–250/hr labour finding below moves cost at The security read and the compute finding does not. The wall-clock seconds for a clone-and-boot of a specific topology remain unpublished, and are now a two-hour engineering exercise rather than a research question.

✔ Gray Swan's corpus numbers

ANSWERED, and the disputed figure does not exist. The Arena About page reports 4M+ attack attempts, 130K+ successful breaks, 13K+ community members, $490K+ rewards distributed, 100+ participants placed in paid roles. The homepage says "over three million" attempts and "15,000+" researchers — date drift on stale marketing copy, not a discrepancy. The "more than one million high-quality, real-world attack trajectories" figure appears in no Gray Swan primary source, on the site, in the Series A release, or on arXiv. Do not cite it.

The derived economics stand and get slightly worse for the crowd: $3.77 per verified break, $0.12 per attempt, $32.67 lifetime per member — against an Arena licence that is "irrevocable, worldwide… for any purpose". See Paying the crowd.


What remains, re-ranked

1. What does a lab pay, per hour, per report or per environment?

Now the only genuinely load-bearing unknown, and the follow-up made it sharper rather than closing it. The contractor-side per-hour question is resolved: $70–90/hr for vulnerability classification, $80–95/hr for senior blue/red analysis, $80–90/hr for GRC and SOC scenario authoring, $200–250/hr for demonstrated CVE authors and top-tier CTF placers doing benchmark construction and expert grading (Mercor). That is a 2.5–3x credential premium and it is a cost input.

No sale price exists at any lab, for any unit. OpenAI pays external testers through "direct payment and/or API credits" with no amount disclosed (OpenAI). Per-hour is the only unit anyone quotes anywhere; there is no per-trajectory, per-report or per-environment price in public. The two nearest analogues are the Hack The Box £39,366 above and DARPA AIxCC's ~$152 per competition task for a fully automated system.

What would settle it. One conversation with someone lab-side who has signed a cyber evaluation contract. Failing that, ask a seller for a quote: Gray Swan's "Request an Evaluation" form on its adversarial-evaluation page, or Ophion Security's published pricing page — the only one of the six firms with public prices, and therefore the closest thing to a public rate for AI-assisted vulnerability triage. Difficulty: low, interview-only. Value: highest remaining.

2. Has any lab ever bought a detection-engineering dataset?

OpenAI's Cyber Blue Team postings state an intent to "create and validate detections" at $293–385K. That proves intent to build the product. It does not prove they buy the data rather than generate it from their own SOC [UNVERIFIED].

Why it matters. Detection engineering is the second-ranked security sub-market on composite, and the only one with a mechanical verifier over data that is legally clean — SigmaHQ rules are DRL-1.1, which explicitly permits selling copies. The whole case rests on a buyer who has been shown to want the capability and never shown to purchase the input. See Detection engineering.

What would settle it. Ask the Cyber Blue Team's hiring manager, or ask SOC Prime whether any lab has licensed Threat Bounty content. Difficulty: low.

3. Who won ARIA's Safeguarded AI Track 2?

Still not announced, and now proven not announced rather than merely not found. ARIA's live funding page still reads "we will announce the successful Creators in the coming months"; the funded-projects page carries no cybersecurity entries. The 31 July 2026 date was a notification date "subject to due diligence and negotiation", with signature expected within four weeks (ARIA).

The contract itself is now fully specified, and it is worth reading as a template: £20m call; Track 1 three-to-six blue teams at £2.5–3.5m each; Track 2 exactly one red team at £2–3m; project start 31 August 2026 to end November 2027; 8-week sprint cycles against 3–5 blue teams simultaneously; procured as a "service contract with ARIA, with deliverables, acceptance criteria, and payment milestones tied to sprint cycles" — commercial terms, not grant terms; >50% of costs and personnel in the UK; compute budgeted at up to £200k per sprint; a three-page proposal (solicitation PDF).

And the call is explicitly rolling — "if we do not identify suitable teams, we will continue to accept and assess proposals on a rolling basis." A non-award is itself a live commercial opening. What would settle it: email ARIA, or FOI the award. Difficulty: very low.

4. What do frontier labs require of a data vendor?

Partially answered, and the earlier page was wrong in one direction. No lab publishes a supplier security standard — OpenAI's Trust Portal lists what OpenAI holds (SOC 2 Type 2, ISO 27001/27017/27018/27701/42001, FedRAMP 20x, PCI DSS) and specifies nothing of vendors; Anthropic's trust centre is JS-gated and anthropic.com/supplier-code-of-conduct 404s.

But the constraints are being imposed elsewhere. Nationality and background checks do flow down through Mercor: the $70–90/hr listing for "a leading AI lab" requires being "currently based in the U.S., Canada, UK, Australia, or New Zealand" and the "ability to pass an enhanced background check." The gating scales with the sensitivity of the data touched, not with seniority — it lands on the listing closest to a lab's internal vulnerability corpus, not on the $200–250 one. And Anthropic's Cyber Verification Program excludes zero-data-retention organisations outright: "Organizations on Zero Data Retention (ZDR) are not currently eligible to participate in the CVP" (Anthropic support). A vendor that promises its buyers ZDR cannot use Anthropic's own models for offensive work — a real operational constraint nobody in this market has written down.

No evidence of any air-gap requirement on a data vendor was found. What would settle the rest: apply for the CVP, which is free and self-serve, and read a lab's DPA and subprocessor list. Difficulty: trivial.

5. Gray Swan's revenue mix

Not disclosed and not derivable, but the structure is now visible. Gray Swan runs two separately quota-carried segments: a Strategic Account Executive – Frontier Labs reporting to an SVP of Frontier Lab Safety — "our frontier lab business is Gray Swan's most strategic… most of our growth comes from deepening engagements that already exist" — and a Senior Account Executive – AI Security at $314–384K OTE selling paid pilots into financial services, healthcare and insurance. Deals in both segments are described as six-figure-plus. Company size is stated in every posting as "approximately 50 people" (Ashby).

Add the £702,730 from UK AISI and one line of the mix is now a fact rather than an inference — and it is a government line, not a lab one. ~50 staff, expansion-led six-figure deals and $40M raised implies revenue in the high single-digit to low double-digit millions, consistent with the low end of the $15–40M estimate for the whole named-vendor segment. What would settle it: investor materials or the next round's coverage. Difficulty: medium.


The conflicts, now resolved

Gray Swan's corpus size — resolved above. The million-trajectory figure is not a Gray Swan number.

Mercor's rate bands — resolved. All four are live listings for four different jobs, read out of the structured job payload rather than rendered copy. The consequence stands: "Mercor cannot reach the exploit-development tier" is wrong and must not appear in any pitch. The still-true version is that Mercor's volume posting is GRC and SOC scenario authoring, and that no generalist demonstrably operates a reproduction-capable triage panel, a range estate, or a contamination-protected held-out set.

The Five Eyes restriction — resolved, and the earlier page was too confident in the other direction. The restriction is real, but it applies to the $70–90/hr listing, not the $200–250/hr one. Do not state it as a blanket rule and do not state its absence as one either.


The structural holes

Things nobody has written down, as distinct from things this research failed to find.

Who funded SRE-Bench's 5,000 expert-hours — the most important commercial question raised by the most important paper in the sub-market research, and completely unanswered.

CVE and NVD licence terms. Both sites are JavaScript-gated and three fetch paths failed. CVE data underpins vulnerability research, appsec and GRC alike, so this is a live legal item, not a curiosity. Same for the CIS Benchmarks' Creative Commons variant — if it is NonCommercial, a commercial cloud-hardening dataset built on it is barred — and for vx-underground and VirusTotal redistribution terms.

SOC Prime publishes no per-rule payout, and cannot. The model is a usage royalty: "the amount of your bounty depends on the popularity of your published rules," rating-based, individuals only, authorship retained (SOC Prime). A royalty has no per-artefact price by construction. Any client offering a fixed, disclosed, per-artefact rate would be differentiated against the only comparable in the market — and should expect the low-quality submission flood that forced SOC Prime's rating system, and gate for it before launch rather than after.

No bug bounty platform publishes a take rate (What a rake can actually be). No cyber-specific exclusivity premium exists. No BIS or EU guidance addresses AI training corpora as controlled technology (Exploits are free, uploading is an export). No offensive-security insurance pricing exists publicly (Nobody prices this risk yet).

No survey of practitioner attitudes to selling training data exists. Every claim here about what the community will accept is drawn from the HackerOne episode. A twenty-person practitioner survey before launch would settle the pricing and consent-design questions at once.

And whether cyber evaluation spend is growing or flat still cannot be shown. The DSIT series gives five points but ends at September 2025, and the publication lag is roughly ten months — too slow to serve as a time series.

The verdict these gaps sit under is The security read; the sequence for closing the top two is Ninety days in security.