The received wisdom, sourced to Anthropic and repeated everywhere, is that sanitizers "perfectly separate real bugs from hallucinations". That is true of the narrow memory-safety case and badly false as a claim about vulnerability-validation artefacts in general. A pre-registered independent reproducibility audit published in August 2026, across a 104-paper consensus corpus of LLM and agent-driven vulnerability workflows, found (arXiv 2608.09567):
- only 59 of 104 (56.7%) papers have a publicly reachable artifact;
- only 10 of 18 (55.6%) sampled artifacts complete their declared workflow — 11 of 18 after environment-only repair;
- 58 of 102 (56.9%) anchor-benchmark cases contain a script-internal CVE identifier that diverges from the declared directory CVE;
- 20 of 30 patched-counterfactual audits still produce the claimed signal on the patched build, and 7 of 19 matched negative controls trigger on benign input;
- the oracle confusion matrix runs at sensitivity 60%, specificity 45%.
The author's conclusion is the sentence to build a company on: "A trigger on the vulnerable build is not evidence of CVE-specific reproduction without a clean patched counterfactual."
A second August 2026 paper reaches the same place from the CTF side. CTF-ABACUS audited 1,435 attempts by six models across 240 challenges, producing 2,870 solve profiles, and found that trace-verified exploits account for only 62–87% of recovered flags — the remainder attributable to direct flag exposure, memorised recall, external lookup, guessing, or unsupported claims (arXiv 2608.26237).
Put together: between 13% and 38% of "solved" challenges were not solved by exploitation, and the automated oracle everyone treats as ground truth is worse than chance on specificity. That converts the tier of cyber data the market treats as a product into a service. Somebody has to build the patched counterfactual and audit the trace. That somebody is an expert, and that is the sellable thing.
What the artefact is
Source or binary in; a novel bug, a working proof of concept, a chained primitive, a severity assessment out. For RL, the exploitation trajectory itself.
But the audit above reorders those by value. The finding is cheap and increasingly automated. The adjudication — reproduce it, diff it against the patched build, confirm the trace actually exercises the flaw, and rate the severity — is expensive, recurring, and has a named buyer with a queue. Anthropic's coordinated-disclosure programme ran 26,153 candidate findings down to 4,576 reviewed by external security firms, and states plainly that "the process of independent human triage and review is the rate limiting step" (Anthropic CVD). Six firms are already absorbing that volume. See What the labs pay for who they are.
A third artefact is emerging and nobody supplies it: stealth. StealthBench extracts 11 hand-verified OPSEC incidents from real bug-bounty and red-team trajectories into 14 dockerised scenarios, graded by a three-model judge panel with majority vote, and finds no model exceeds a 54% safe success rate — solved and undetected (arXiv 2607.26314).
Is there a verifier
Not the one the market thinks it has. Grade it in three tiers and price them differently:
| Tier | Verifier | Reality |
|---|---|---|
| Memory-safety crash under a sanitizer | Genuinely deterministic | And therefore commoditised — Bugcrowd's OSS-derived environments already cover it at volume |
| Artifact-embedded oracle asserting "CVE-X reproduced" | 60% sensitivity, 45% specificity | Fails on two thirds of patched counterfactuals; 56.9% of anchor cases carry a mismatched CVE ID |
| Logic bugs, business-logic flaws, exploitability, blast radius | Expert judgement only | No oracle exists and none is coming |
Defense scores 4 on the strength of the bottom two rows. If the oracle were perfect, this would be a data business with a decaying price. It is not perfect, and the gap between "the tool fired" and "the bug is real" is a recurring human cost that does not go away as models improve — it grows with model output volume. The one shape that survives is the full form of this argument.
Proof scores 3, not higher. Being visibly the best adjudicator requires publishing an audit methodology and a track record, which takes a cycle, and six firms are already doing the work with lab logos attached.
Is anyone buying
Budget scores 5, uncontested. These are the highest security bands found anywhere in this research.
| Buyer | Role | Band |
|---|---|---|
| Anthropic | Lead, Frontier Red Team (Cyber) | $485,000 – $755,000 |
| OpenAI | Offensive Security Engineer, Agent Products — "principal-level… deep, hands-on penetration testing of OpenAI's agent-powered products" | $347K – $490K |
| OpenAI | Researcher, Frontier Cybersecurity Risks (Preparedness) | $295K – $445K |
| Anthropic | Research Engineer, Cybersecurity RL | $300,000 – $405,000 |
| Anthropic | Head of Vulnerability Disclosure & Security Community | $330,000 – $395,000 |
| Scale AI | Strategic Projects Lead, Red Team | not disclosed |
Read the fourth row twice. A "Research Engineer, Cybersecurity RL" role existing at all is direct confirmation that Anthropic is building cyber RL environments in-house. It is simultaneously the best demand signal in the dossier and the clearest competitive threat to an external environment vendor. Room scores 2 largely because of it, and because six external firms plus Bugcrowd plus Mercor's own $200–250/hr bench are already in the space.
What the expert costs
Cost scores 2, and the reason is not the hourly rate. It is the outside option.
Crowdfense runs a $30 million acquisition fund, up from $10M in 2017, and publishes a rate card offering $10,000 to $7 million per capability (Crowdfense):
| Capability | Price |
|---|---|
| iOS zero-click full chain | $5M – $7M |
| Android zero-click full chain (WhatsApp, RCS) | $5M |
| Safari one-click full chain | $2.5M – $3.5M |
| Chrome one-click full chain | $2M – $3M |
| Windows zero-click full chain | $1M |
| Chrome / Safari RCE without sandbox escape | $500k |
| Firefox full chain | $300k |
| WinRAR RCE | $100k |
ZipRecruiter's $113,102/yr, $54.38/hr for a vulnerability researcher is a nonsense figure against that. The real number is Mercor's $200–250/hr band for Cybersecurity Research Expert — Offensive Security & Vulnerability Research, which requires multiple of: top-ranked CTF team membership, CVE discovery, or publicly recognised bug bounty work, and whose task is explicitly benchmark construction (listing).
The arithmetic that matters: a person who can produce a Firefox chain has a $300,000 outside option. Paying them $250/hr buys 1,200 hours for the price of one bug they might not find — which is why the tier exists. But it also means no rate card buys the top decile away from the exploit market. You are only ever buying spare capacity. Offensive supply and Paying the crowd work this through.
Getting to them
Reach scores 4. The registers are real and named, and the top of each is public.
- CTFtime: 38,575 registered teams. Player count is not published, but the leaderboard names the people the Mercor listing is asking for.
- HackerOne: roughly 50,000 researchers who have ever earned a bounty, with the top 1% earning more than the bottom 90% combined and 30 hackers above $1M lifetime. The distribution is the point: you want a few hundred names, not fifty thousand.
- Synack Red Team: 1,500+ — the only fully vetted population, and the best available proxy for professional-grade reliability.
- Pwn2Own / ZDI remains the public leaderboard for chain-capable researchers (ZDI blog); 2026 prize totals could not be established.
- OSCP and OSCE holder counts have never been published by OffSec.
Where the benchmarks sit
| Benchmark | Frontier score (Claude Opus 5 unless noted) |
|---|---|
| Firefox 147 exploitation | 131 full exploits / 250 trials (52.4%); Mythos 5: 221/250 (88.4%) |
| OSS-Fuzz | 79.4% non-zero scores; Mythos 5: 80% >0, 13 complete exploits |
| ExploitBench (V8) | 10.14 mean flags, 99 full ACE exploits; Mythos 5: 10.80 / 132 |
| Irregular CyScenarioBench | 33.7% challenge completion; Mythos 5: 47.0% |
| UK AISI cyber ranges | "The Last Ones": 8/10 end-to-end; "Doing Life": step 22 of 23 |
| StealthBench | ≤54% safe success rate, all models |
| CTF-ABACUS trace audit | only 62–87% of recovered flags are trace-verified exploits |
All Opus 5 figures from the Claude Opus 5 System Card. The offensive curve is the one saturating fastest in the whole dossier — see The measurement gap. That is precisely why the durable product is the audit and not the discovery.
What would kill it
Your suppliers get a better offer, permanently. They already have one. A seven-figure rate card is not something a data business competes with; it is something it schedules around.
The labs internalise triage. Anthropic already runs a Cybersecurity RL engineer and a Head of Vulnerability Disclosure. If human review stops being the rate-limiting step — through better oracles or cheaper models — the adjudication line contracts.
Someone publishes the counterfactual harness. The audit that exposed the 45% specificity also describes how to do it properly. A well-engineered open patched-counterfactual pipeline would automate the cheap half of adjudication and leave only logic bugs, which is a smaller business than this page describes.
Where the record is thin
CVE and NVD licence terms could not be read — both sites are JavaScript-gated, and the CVE Program Terms of Use content was unretrievable. Given that CVE data underpins this sub-market, Application and cloud security and Governance, risk and compliance, that is a live legal item, not a footnote.
Pwn2Own 2026 prize totals, OSCP/OSCE holder counts and HackerOne's 2026 platform statistics were all attempted and none retrieved. And no external adjudication contract value is public anywhere — the closest proxies are UK AISI's payments, set out in What the labs pay and Government buyers.