Every question about whether a cyber expert-data company is a product business or a services business collapses into one technical property: can the environment tell, without a human, whether the model actually did the thing?
The evidence on this is unusually clean, and it comes from the buyer. Anthropic states that for memory-corruption bugs, "tools like Address Sanitizer perfectly separate real bugs from hallucinations," and reports zero false positives testing Opus 4.6 against Firefox 112, alongside 181 successful working exploits and 29 register-control achievements against the Firefox JavaScript engine (Anthropic). In the same document, on logic vulnerabilities, they say the opposite: "we do lose the ability to (near-)perfectly validate the correctness of any bugs Mythos Preview reports."
Those two sentences are the whole strategy question. A perfect oracle means an artefact a buyer can run ten thousand times unattended — a licensed corpus, sold once, reusable, RL-trainable. No oracle means a permanent human panel attached to the artefact, priced by the hour, scaling with headcount, and never leaving.
Oracle quality is not a technical detail underneath the product. It is the product's business model, and it is legible from the outside before you write a line of code.
Where the oracle is perfect, someone else already owns it
| Oracle quality | Example | Sellable as | Margin shape |
|---|---|---|---|
| Perfect — sanitizer crash, flag or canary | memory corruption, CTF-style tasks | Product: licensed corpus, reusable, RL-trainable | High — and already commoditised |
| Partial — multi-step range with per-step checks | AISI's "The Last Ones", CyScenarioBench | Product plus maintenance: needs re-authoring as models saturate it | Medium, with recurring re-authoring revenue |
| None — logic bugs, severity, business-logic auth, attribution | Anthropic's logic-bug pipeline; threat-intel attribution | Service: retained expert panel, per report or per hour | Labour margin, scales with headcount |
The top row looks like the attractive one. It is the one to avoid.
Bugcrowd launched RL Environments for frontier AI labs on 21 May 2026, built on its November 2025 acquisition of Mayhem Security — symbolic execution and fuzzing descended from DARPA's Cyber Grand Challenge. The product supplies "hundreds of thousands of training environments, each built from open-source software with real source code and verifiable outcomes," in which agents locate a bug, trigger it, assess exploitability and produce a fix, with objective scoring at each step. Frontier labs and LLM providers are already using it. Bugcrowd states explicitly that "no customer data or work from its researcher community is used in the environments." CAIO David Brumley: "You cannot train a model to be good at security by showing it what security looks like, you have to give it real problems to solve" (SiliconANGLE).
Read that against the ARIA and AISI numbers below. A fuzzing company with a corpus of open-source repositories generates perfect-oracle environments by the hundred thousand at near-zero marginal cost. A specialist authoring them by hand at $80–90 an hour is competing with a build script.
The perfect-oracle tier is where the scaling is, which is exactly why it is already occupied. AIxCC reinforces the point from the other direction: a fully autonomous 143-hour competition across 48 challenge projects, validated on "crash stack traces, sanitizer signatures, and automated execution," with all seven finalist systems open-sourced (AIxCC SoK). The free corpus and the commodity corpus both sit in the same tier.
Where there is no oracle, the demand is documented in headcount
Anthropic's coordinated-disclosure programme is the clearest picture anyone has published of what happens when the automated verdict disappears: 26,153 candidate findings → 5,008 advanced to independent review → 4,576 confirmed valid (91.4% true-positive) → 2,300 disclosed across 392 open-source projects. Triage is performed by "one of six external security research firms" who "reproduce each issue, assess whether it is a real bug." Severity agreement with Claude ran 85.2% exact, 97.1% within one band (Anthropic CVD dashboard); on a separate 198-report sample, 89% exact and 98% within one severity level (Anthropic).
Five thousand human-reviewed reports, and one lab needed six separate firms to absorb the volume. That is the highest-confidence demand signal anywhere in this dossier, and it exists precisely because the oracle failed. The labs as buyers traces the rest of the buying behaviour.
Even the vendor with the strongest product story keeps humans in the loop at the open-ended tier: Irregular's FrontierCyber outcomes are determined by "automated checks and expert review" (FrontierCyber).
The infrastructure that breaks the cheap-container economics
The general RL-environment market runs on containers. SemiAnalysis describes the standard pattern as "dockerized containers, MCP servers exposing tool definitions, prompt/setup conditions, success criteria, reward signals, and telemetry capture" (SemiAnalysis). That is what makes a $20,000 website replica a $20,000 website replica.
UK AISI explicitly rejected containers for cyber. Their ranges run nested virtualisation on "beefy, bare metal servers" — hypervisors with VMs nested inside — because it gives "an incredibly deep amount of control over the network" and lets them host Windows and real-time operating systems rather than Linux alone. The tools named are ProxMox and Vagrant (FAR.AI, Mahmoud Ghanem). Irregular goes further still: FrontierCyber runs against physical devices — phones and routers running real applications, plus image libraries inside file-processing services and browsers rendering attacker-controlled pages (FrontierCyber).
Anthropic's own vulnerability-discovery environments are containers — but they are single-project, perfect-oracle, memory-safety environments, which is the tier where containers work.
So the cost structure splits along the same line as the oracle. Perfect-oracle environments are cheap containers that someone else mass-produces. The environments a specialist would actually sell — multi-host, non-Linux, network-realistic, physical-device — need an estate, a hypervisor fleet and an operations posture.
This is the live diligence item, and it decides whether these artefacts are RL-usable or evaluation-only.
RL training needs thousands of resets; Kimi's infrastructure is reported to instantiate 10,000+ environment instances at once (SemiAnalysis). On bare-metal nested virtualisation, a reset is a snapshot restore of an entire VM topology, materially more expensive than a container teardown. I could not find a published figure for reset time or cost on a multi-host cyber range anywhere [COULD NOT ESTABLISH].
Hack The Box's answer to the same problem is to sidestep it: continuous refresh rather than reset — weekly new target releases plus "multiple overlapping evaluations across scenario variations to prevent overfitting" (HTB AI Range). That is an evaluation product's answer, not a training product's answer. Until someone publishes reset economics, treat any claim that a bare-metal range is an RL asset rather than an eval asset as unproven.
What one costs to build
The only two ranges with published human-effort figures are both AISI subcontracts, and both were built by outside firms (AISI):
| Range | Builder | Shape | Human expert solve time |
|---|---|---|---|
| The Last Ones | SpecterOps | 32 steps, 4 subnets, ~20 hosts; recon, credential theft, lateral movement across multiple AD forests, CI/CD supply-chain pivot, exfiltration | ~20 hours |
| Cooling Tower | Hack The Box | 7-step ICS attack on a simulated power plant, including reverse-engineering a proprietary control protocol and its cryptographic authentication | ~15 hours |
Those are solve times, not build times. A common heuristic in evaluation work is 10–50x solve-time to author, QA and harden; at 20 solve-hours that implies 200–1,000 expert-hours, or $16,000–$90,000 of direct expert labour at Mercor's $80–90/hr cyber-expert band [UNVERIFIED — the multiplier is an assumption, not a sourced figure]. Add AIxCC's evidence that ground-truth annotation alone on a 63-vulnerability corpus consumed two person-weeks of manual verification by two independent reviewers plus 8,906 CPU-hours of reference fuzzing (AIxCC SoK), and one serious multi-host range plausibly lands at $100k–$400k fully loaded — consistent with Epoch's $300k figure for a complex product replica. An academic build gives a third data point: six senior experts (5+ years each) produced 266 instances across 8 ranges and 156 hosts (AgentCyberRange).
AISI's own range designer names the binding constraint, and it is not money: "people with these skills are busy defending actual systems."
The prices those costs are sold into are in What you can actually sell, and the total pot they add up to is in Sizing the cyber pot.
Exclusivity: 4–5x, and it buys the wrong thing
The only hard number is Epoch's: "Exclusive deals are roughly 4-5x more expensive than non-exclusive ones," given independently by two RL-environment founders — with Epoch noting the article "does not define what 'exclusivity' operationally means" (Epoch). No cyber-specific exclusivity multiple exists [COULD NOT ESTABLISH].
What a lab is buying with that premium, on the surrounding evidence, is mostly not capability advantage. Irregular keeps its evaluation set "private to avoid contamination" and advertises "proprietary in-house cyber challenges (avoiding training data contamination)" (CyScenarioBench). A shared evaluation environment leaks into rivals' training data and stops measuring anything at all. For an evaluation artefact, exclusivity is not a premium feature, it is the entire product — which is a much stronger argument for a higher cyber multiple than "our gym is better than their gym," though it remains inference [UNVERIFIED].
There is a third thing bought, visible in the Anthropic–Irregular incident, where evaluation machines had live internet access despite prompts telling Claude they did not, and Claude reached real third-party systems (Anthropic). Exclusivity also buys control over who else knows how the range behaves. A marketplace cannot carry that liability; see The generalists in cyber.
The best-documented contract structure in an adjacent vertical is FrontierMath: OpenAI retains ownership of the questions, receives statements and solutions for 300 problems and statements only for a 50-problem hold-out, and Epoch "cannot share questions and answers without written permission" while retaining the right to run and publish evaluations (Epoch). Licence plus hold-out, not assignment. That is the template — and the same hold-out logic is what makes Publishing the benchmark work at all.
Sell the environment as the wedge. The judgement attached to it is the business, and The measurement gap shows where the judgement is scarcest.