Miju Labs

The security dossier

The security read

Build, still — but on worse terms than the first reading. The six firms are named, one of them already sells this exact product to at least two labs, the first contract values in the market's history are now public from UK transparency data, and the elite labour tier costs three times what the earlier estimate assumed.

medium confidence14 minupdated 2026-08-30verdict · scoring · oracle · logic bugs · frontier labs · contract values

Build, but not the company the screen described. The defensible business in security is not a cyber data company, and the choice is not between offensive and defensive. It is a human validation service for vulnerabilities that have no automated oracle, with held-out evaluation content falling out of it as a by-product and an environment corpus as the marketing artefact. Everything else in this dossier — the trajectories, the memory-safety gyms, the public benchmarks, the crowd — is already commoditised, already owned, or unpriced. That single surviving shape is argued in full at The one shape that survives. This page is why nothing else is left standing.

This page supersedes the earlier reading

A follow-up pass closed three of the gaps this verdict was written around, and all three moved the numbers. The six external security research firms are named. Contract values now exist in public, from UK government transparency data, retiring the claim that none did. And the elite offensive labour tier resolves at $200–250/hr, roughly tripling the per-environment cost model. Two of the six scores move as a result — cost from 3 to 2 and room from 3 to 2, taking the total from 22 to 20. Where an earlier sentence conflicts with what is below, the sentence below is the current one.

What the deep look changed

The shallow screen ranked Defensive security third overall in the atlas on two claims: the widest published benchmark gap of any domain, and no incumbent. It ranked Offensive security first on evidence and last on availability. Three findings from the deep research move those numbers, and they do not all move them the same way.

A direct incumbent exists, and it is cheap. Loginsoft markets "Security Data for AI Training" — "curated, labeled, and synthetic cybersecurity datasets" with "expert labeling and ground-truth validation" — plus a separate AI Model Validation line building "evaluation datasets that reflect real operating conditions" with "human-in-the-loop review" (Loginsoft; model validation). Founded 2005, HQ Chantilly VA, development centre Hyderabad, 201–500 people, billing $25–49/hour [WEAK — third-party directory] (Enosis).

That is disconfirming, on two axes at once. The screen's room: 5 was wrong — the category was not empty, it was unnoticed. And the cost assumption was wrong: expert security labelling is not inherently expensive, because a firm with two hundred people in Hyderabad will do it for a third of Mercor's posted rate. What it does not disconfirm is demand. Loginsoft has no venture backing, no benchmark, no publication record, no lab logo, and names no client for the AI data line on its own site. It is the price floor, not the market. See The defensive vendors.

Mercor already recruits at the exploit-development tier, and the rate bands now resolve into a clean credential premium. Read out of the live listings' structured payload rather than the rendered copy, four bands describe four different jobs and none contradicts another: $70–90/hr for a two-plus-year engineer classifying vulnerabilities for "a leading AI lab"; $80–95/hr for a senior blue/red analyst; $80–90/hr for a SOC and GRC practitioner authoring scenarios and rubrics against NIST CSF, SOC 2, ISO 27001 and NIS2; and $200–250/hr for someone who has actually shipped CVEs or placed on a top-ranked CTF team, on a posting whose task is to "evaluate AI-generated analyses of complex cybersecurity scenarios, exploits, and vulnerability reports" and "contribute to benchmark development" (Mercor).

That is a 2.5–3x premium for demonstrated elite offensive credentials over generalist senior security practice, and the top band is the addressable price point for this business.

The cost model roughly triples

The earlier sizing used $70–90/hr and produced $15K–$70K of raw expert labour per premium environment. Built by the credentialed tier at $200–250/hr, the same environment is $40K–$200K. Cross-check it against the one public per-environment number in existence: UK AISI paid Hack The Box £39,366 for the "Cooling Tower" ICS range, which takes roughly fifteen human-hours to solve; at the 10–40x build-to-solve ratio in The oracle decides everything that is 150–600 build hours, an implied blended £66–260/hour. The elite Mercor band sits inside that bracket.

Two consequences. The "generalists cannot reach that supply" argument is dead and should not be used. And the elite tier is priced above what the vendors pay their own staff — a Gray Swan Red Team Engineer tops out at $185,000, about $89/hr fully utilised (Ashby) — which is the clearest available signal that scarce credentialed offensive talent, not tooling, is the binding constraint here. What Mercor does not run is the infrastructure: a reproduction-capable triage panel, multi-host ranges, a contamination-protected held-out set. That gap survives, narrower than the screen implied. The generalists in cyber reconciles all six bands.

Both bounty platforms examined the licensable-trajectory business and declined. HackerOne's terms leave IP with researchers, its CTO stated it does not train LLMs on submissions, and its CEO repeated the position in February 2026 — "You are not inputs to our models" (The Register). Bugcrowd, sitting on 500,000+ researchers, launched RL Environments for frontier labs on 21 May 2026 built entirely from open-source CVEs, stating flatly that "no customer data or security researchers are used at any stage of the training process" (Bugcrowd).

This one is genuinely ambiguous and should not be spun. The bearish reading is that the two parties with every advantage looked at the problem and routed around it, which is evidence the corpus cannot be assembled at a price anyone will pay. The bullish reading is that neither declined on economics: HackerOne declined on contract, Bugcrowd on product shape — its environments need reproducible oracles that crowd submissions do not have, and Mayhem gave it a cheaper route. Neither tested paid, consented, purpose-built hourly production. Both readings fit the same facts, and this is the largest single uncertainty in the thesis. The bounty platforms sets both out at length.

The case for, which is stronger than the screen knew

A frontier lab stopped a training run on cyber grounds. On 7 August 2026 OpenAI announced that its upcoming model Astra had reached the Critical cybersecurity capability level under its own Preparedness Framework (OpenAI). Per Axios it halted two weeks of deployment-focused RL training, put its largest planned frontier RL run on hold, left "a significant number of Astra and cyber-related research workloads" paused, and began rewriting the framework — while committing to give "recommended security controls to third-party testing partners for running higher risk evaluations" (Axios). Cyber evaluation is not a discretionary line at OpenAI in H2 2026. It is the gating item on a flagship model.

The other anchor buyer published the sales pitch about itself. The Claude Opus 4.6 system card states that the model "has saturated all of our current cyber evaluations," at ~100% on Cybench (pass@30) and 66% on CyberGym (pass@1), and that "[t]he saturation of our evaluation infrastructure means we can no longer use current benchmarks to track capability progression" (system card). Anthropic's response was to invest in harder evaluations — with no formal RSP cyber threshold at all. By Opus 5, CyberGym is retired.

Six external security research firms are already being paid for exactly this service, and they are now named. Anthropic's coordinated-disclosure programme ran 26,153 candidate findings → 5,008 advanced to independent review → 4,576 reviewed by external firms at a 91.4% true-positive rate → 2,300 disclosed across 392 open-source projects, with 462 identifiers issued (177 CVE, 285 GHSA). Severity agreement with Claude runs 85.2% exact and 97.1% within one band (Anthropic CVD). A separate 198-report sample reviewed by "professional human security contractors" agreed 89% exactly and 98% within one level (Mythos Preview).

The list sits one click deeper than the announcements, on the "About this dashboard" page of the same site: Ada Logics, Anvil, Calif.io, Doyensec, Ophion Security and Trail of Bits (Anthropic CVD — about). Two are observable in the wild: wolfSSL's advisory list carries eight advisories credited "Thanks to Calif.io in collaboration with Claude and Anthropic Research" and three to "Max at Trail of Bits" (wolfSSL).

Five thousand human-reviewed reports at one lab, six firms to absorb the volume, and Anthropic's own statement that "the process of independent human triage and review is the rate limiting step". That is a capacity-constrained, headcount-scaling, recurring service with a named buyer and, now, a named supplier list.

Calif.io already sells this exact product to at least two labs

The most commercially significant of the six is Calif.io, founded by Thai Duong of BEAST and CRIME. Its site states: "We partner with frontier labs and stay with our customers for the long haul. Half our work is securing AI." One published service line is "Capability evaluation — We benchmark what a model can actually do offensively, measured against real targets rather than capture-the-flag toys, so the claims you publish are ones you can defend." Its logo wall includes Google DeepMind; its testimonials include Anthropic's Deputy CISO and an Anthropic security engineer (calif.io).

That is the wedge, productised, sold to at least two frontier labs, by a firm most market maps do not list. It is the single most disconfirming fact in the follow-up, and it is why room drops to 2.

The other five are a mix rather than a wall of consultancies: Ada Logics, a long-standing OSS-Fuzz and OpenSSF audit contractor with an AI/LLM line; Doyensec, an offensive boutique with a published LLM practice; Anvil Secure, an employee-owned US pentest firm with a named AI Security Testing service; Trail of Bits, the incumbent and also a UK AISI supplier; and Ophion Security, a product company that is the only one of the six with a published pricing page — the closest thing to a public price for AI-assisted vulnerability triage anywhere.

The defensive capability gap is vendor-confirmed and wide. CrowdStrike and Meta's CyberSOCEval puts frontier models at 23–34% on malware analysis and 43–53% on threat-intelligence reasoning, against random baselines of 0.63% and 1.7%, and says plainly that "current LLMs are far from saturating our evaluations" (arXiv 2509.20166). Microsoft's ExCyTIn-Bench tops at 56.2%; CTI-REALM at 0.637, and 0.282 on cloud multi-step tasks (arXiv 2603.13517).

And reasoning models do not close it, which is the load-bearing part. CyberSOCEval reports that test-time reasoning models "do not achieve the boost they do in areas like coding and math, suggesting that these models have not been trained to reason about cybersecurity analysis." CTI-REALM, built by a different company on a different task, independently found medium reasoning effort beating both high and low, and dedicated reasoning models underperforming general-purpose ones.

Why that last finding is the pitch

If more inference-time compute fixed defensive security, the gap would close on its own schedule and no purchase would be needed. Two hyperscalers independently finding that it does not is the best available evidence that the bottleneck is training data rather than scale — and specifically that nobody has ever trained a model on a corpus of practitioners reasoning through defensive work, because no such corpus exists. Everything in The measurement gap points the same way: multiple-choice security knowledge is saturated, doing the job is not close.

Two structural tailwinds sit under all of it. DARPA converted $9.9M of public money into permanently free, open-source cyber reasoning systems and an open challenge corpus at AIxCC (DARPA), contaminating the free corpus for benchmarking and raising the price of private held-out content. And the Frontier Model Forum, covering thresholds from six labs, concedes in writing that third-party "evaluation maturity remains uneven" and that the ecosystem needs "standardized cyber evaluation frameworks, shared red-teaming methodologies, and common benchmarks" (FMF).

The case against, at the same volume

Contract values now exist — and they are smaller than the enthusiasm. The earlier version of this page opened this column with "no contract value exists in public anywhere." That is no longer true. The values were not on Contracts Finder or Find a Tender; they were in DSIT's monthly transparency publication of spend over £25,000, line by line, with a dedicated "Safety – Cyber" cost centre (DSIT).

SupplierPaid by UK AISIWhen
Gray Swan Security Inc£702,730 (£638,345 + £64,385, same day)31 Mar 2025
Irregular (as Pattern Labs Tech Inc)£459,000 across three paymentsMar–Jul 2025
Trail of Bits Inc£245,967 across three paymentsJan–May 2024
Crystal Peak Security LLC£116,66615 Mar 2024
Hack The Box Ltd£39,3669 Jul 2025

Three readings, and they do not point the same way. The market is real and priced: Gray Swan billed one non-lab customer £702,730 in a day, and Irregular — 35 people, $80M raised — books roughly £460k a year from one government buyer. Government is a genuine second buyer, and unusually it names its suppliers in public blog posts and publishes what it paid them. And the absolute numbers are modest: £702,730 is the largest single line in this market's entire public record. The band in Sizing the cyber pot ($25M–$120M/yr, most likely $40M–$80M) survives contact with these figures rather than being displaced by them, and would still be a rounding error at Mercor.

What is still missing is any lab-side value. OpenAI pays external testers through "direct payment and/or API credits" and discloses no amount (OpenAI). No per-report, per-trajectory or per-environment price exists at any lab, and SpecterOps' AISI value could not be found — it is in none of the twenty-one DSIT files from January 2024 to September 2025, because the series ends there. AISI also redacts some suppliers: rows marked "Supplier name withheld" total roughly £1.79m.

One incumbent is load-bearing and would be hard to displace. Irregular is credited in the cyber sections of system cards at OpenAI, Anthropic, Meta and Google DeepMind, on roughly 35 people — "not one vendor with a large market share, but one vendor whose methodology, benchmarks and infrastructure are load-bearing for multiple competitors' safety claims at once" (BERI). Its hold is the rarest kind in the data business: the evaluation set "remains private to avoid contamination," so a customer who takes it in-house destroys it. That is not displaced by being better. See Irregular, in full.

The crowd tier is priced at four dollars. Gray Swan AI has distributed $490,000+ across 130,000+ successful breaks — roughly $3.77 a unit — and the Arena terms grant "Gray Swan AI and its partners an irrevocable, worldwide license to use and share the submission for any purpose" (Arena About). Perpetual, sublicensable, no field-of-use limit. Nobody paying $85–250/hr competes with that on price. Paying the crowd explains why they are different products; it does not make the cost comparison go away.

The perfect-oracle tier is already a commodity. Bugcrowd ships "hundreds of thousands of training environments" from open-source software with verifiable outcomes, at what is effectively zero marginal cost per environment. A specialist hand-authoring memory-safety gyms at $80–90/hr is competing with a build script. The oracle decides everything states the rule: where the sanitizer gives a verdict, do not plan to sell.

And the labs are hiring the function in-house. Anthropic's Lead, Frontier Red Team (Cyber) posts at $485,000–$755,000, its Research Engineer, Cybersecurity RL at $300,000–$405,000, OpenAI's Data Scientist, Cybersecurity at $263,000–$515,000 — roughly 80 cyber-titled open roles across the two anchor accounts (The labs as buyers). In-housing is what a buyer does before a market exists, and also what it does instead of creating one.

Two smaller things belong in the same column. Mercor's answer to a demonstrated environment business is acquisition — Sepal AI in February 2026, Deeptune in July, four months after Deeptune's $43M Series A. And citations do not convert alone: METR is the most-cited evaluation organisation in the field and "has not accepted funding from AI companies" (METR). See Publishing the benchmark.

The six scores

AxisScreen: offensiveScreen: defensiveDeep lookAfter follow-up
proof2544
budget5355
defense3433
cost4432
reach5444
room2532
total21252220

proof: 4 — held, and it was the closest call. Two new facts pull opposite ways and cancel. In favour: UK AISI names its suppliers in public and publishes what it paid them, a route to a checkable reference that no frontier lab offers. Against: the artefact this page called "genuinely unclaimed" is adjacent to a service line Calif.io already sells. What remains strictly unclaimed is narrower than before — no public benchmark measures whether a model's severity judgement on a logic bug matches an expert's, and Anthropic's agreement numbers (85.2% exact, 97.1% within one band) are still the only scoreboard in existence.

budget: 5 — held, and now earned rather than inferred. Defensive was never a separate market: it is the same two buyers and the same evaluation budget, one of which paused a frontier run over it. The follow-up adds arithmetic where there was inference — five named suppliers with pound values against a government cost centre labelled "Safety – Cyber" — plus a new named lab-side buyer for the adjacent RL-environment product: Bugcrowd's own job advertisement names Anthropic, OpenAI and Cohere as frontier labs using environments it builds (Greenhouse). Cohere appears nowhere else in this dossier's buyer map.

defense: 3 — held, deliberately, and the composition of the argument has changed. The bullish half got stronger: the elite offensive tier is priced above what the vendors pay their own staff, which is the clearest possible evidence that the scarce input is genuinely scarce and stays scarce. The bearish half got stronger by the same amount: six named firms, four labs, Mercor, Handshake and a government buyer are all bidding for the same few hundred people, and Anthropic could still ship a severity grader trained on its own 5,008 reviewed reports. A confirmed-scarce input you must outbid five funded competitors for is a 3, not a 4.

cost: 2 — down from 3. The labour basis was wrong by a factor of three: the earlier model used $70–90/hr, and the credentialed tier that can actually do this work is $200–250/hr, taking a premium expert-authored environment from $15K–$70K to $40K–$200K. Everything else that pushed this down still applies — CTI-HAL's human-annotated corpus took two annotators, 81 reports, eight weeks (arXiv 2504.05866); SOC 2 Type II or ISO 27001 is a gate rather than a differentiator; a hold-out set is inventory you build and never sell. The one thing that got cheaper turns out not to matter: AISI's ranges are Proxmox VE virtual machines under an open-source Inspect plugin, where a concurrency slot costs on the order of $0.53/hour and a realistic multi-host range $2–8/hour, so a four-day, 122-run campaign is a few thousand dollars of compute. Infrastructure is not the cost driver here; expert build time is — which is exactly why tripling the labour rate moves the score.

reach: 4 — held, untouched by the follow-up. The channels are named, addressable and cheap: the Critical Thinking podcast, tl;dr sec at 90,000+, Detection Engineering Weekly at 5,200+, SigmaHQ's contributor register, CTFtime's 38,575 teams. Conversion is still the problem rather than contact — Gray Swan turned 15,000 signups into 100+ placeable professionals, 0.7%, and the largest single channel is now operated by an OpenAI employee. Nothing found since changes either half. See Offensive supply and Defensive supply.

room: 2 — down from 3, and this is the substantive downgrade. The earlier version held at 3 on the grounds that "none of those six is a productised specialist in logic-bug validation." That is falsified. The flip condition this page set — "if the six firms turn out to be three consultancies on retainers a new entrant cannot bid against, room drops to 2" — fired in an unexpected direction: not procurement lock-in, but a small, credible, founder-led competitor already inside the position. What keeps this off a 1 is that the six are mostly small specialist firms rather than framework incumbents, that Anthropic needing six of them and calling human triage "the rate limiting step" is the signature of a capacity-constrained market, and that none of them publishes an inter-rater calibration benchmark.

20, against 22 before the follow-up and the screen's 25 and 21. Every closed gap made the market more real and the position more crowded at the same time. That is the normal shape of good news in a market this small.

When the answer flips

Two of the four conditions on each side have already fired, which is why this section is shorter than it was.

To no. If Calif.io raises, hires, or is named in a system card as a capability-evaluation supplier, room goes to 1 and the correct move is to sell into the six rather than become a seventh. If a lab quotes a rate for logic-bug validation at or below Loginsoft's $25–49/hr, the labour arbitrage is gone and this is a Hyderabad business, not a European one. If Anthropic ships a severity grader trained on its own 5,008 reviewed reports and the human panel shrinks rather than grows, the recurring-demand argument dies at its source. And if the lab-side price question stays unanswered through commercial discovery — the only genuinely load-bearing unknown left at What the dossier could not establish — the numerator does not exist and nobody should raise money against it.

To a much louder yes. If a lab states a per-report or per-hour price above roughly $150/hr for validated logic-bug review, the unit economics close on public information for the first time. If ARIA's Safeguarded AI Track 2 award — a single red team at approximately £2–3m over fifteen months, procured as a service contract with milestones tied to sprint cycles, still unannounced and explicitly still open on a rolling basis (ARIA) — goes to a security firm rather than an existing lab evaluation vendor, government is confirmed as a broad second buyer rather than a one-off. And if OpenAI's rewritten Preparedness Framework names third-party validation capacity as a gating requirement, the budget stops being discretionary at both anchor accounts simultaneously.

The DSIT figures have already half-fired the government condition: five named suppliers, five published values, one buyer. That is more than this market had a week ago and much less than a business plan needs.

The one shape that survives all of this is at The one shape that survives. The sequence for testing it in ninety days is at Ninety days in security — and it now starts from a named supplier list rather than from finding one. What is still unestablished, re-ranked around what the follow-up closed, is at What the dossier could not establish.