Start with the honest version of the price table, because the interesting thing about it is what is missing.
There is no public per-environment price for a cyber range anywhere. Not one. Every cyber-specific number below is one of three things: a government contract envelope, the cost of running an evaluation, or a general-RL-environment price applied to cyber by analogy. Anyone quoting a per-range price is quoting a general-market comparable.
| Artefact | Figure | What kind of number this is |
|---|---|---|
| Simple web/UI replica environment | ~$20,000 per website | General RL, not cyber (SemiAnalysis; Epoch) |
| Complex product replica (a Slack clone) | ~$300,000 | General RL, one interviewee (Epoch) |
| Per-task price, typical band | $200–$2,000 per task | General RL, consensus across vendors (Epoch) |
| Per-task, top of market (complex SWE) | up to $20,000, described as rare | General RL (Epoch) |
| Typical vendor–lab contract | six to seven figures per quarter; $300k–$500k at the neolab end | General RL, founder interviews (Epoch) |
| Exclusivity premium | 4–5x the non-exclusive price | General RL, two founders independently (Epoch) |
| Open-source environment bounty | $1,000–$5,000+ per environment | Prime Intellect programme (X) |
| ARIA red-team / evaluation partner | ~£2–3m over ~15 months, one partner | Cyber — but a contract envelope, not an artefact price (ARIA) |
| Cost of running a cyber eval task | AISI at 50M tokens: ~$10 average per run, <$60 max; Irregular: <$1 easy, ~$20 medium/hard, ~$100 hardest | Cyber — but a run cost, not a sale price (Irregular) |
| Expected cost per successful solve | Crypto-misuse challenge $224/success; lateral-movement $965/success | Irregular's own metric (Irregular) |
Solving one AISI advanced task (rust_vm) | $1.73 in API cost (GPT-5.5) | (AISI) |
| Mercor cyber expert / pentest task | $80–90/hr; $1,750–2,150 per completed task | Input cost, not sale price (Mercor) |
No published per-environment price for a cyber range [COULD NOT ESTABLISH], and no disclosed pricing from Irregular, Gray Swan or Hack The Box's AI Range — none of the 38 tracked RL-environment vendors publishes pricing at all (RL List). The AISI subcontract values to SpecterOps, Hack The Box and Crystal Peak Security are not on Find-a-Tender or Contracts Finder, consistent with being sub-threshold or under an existing framework.
Price discovery here happens one negotiation at a time, and the first serious quote a specialist issues sets a comparable for itself.
The units, one at a time
Attack trajectories. One adversarial interaction sequence, labelled success or failure, usually with a judged harm category. Gray Swan's public aggregates are the only real shape available: 4M+ attempts, 130K+ successful breaks, $490–500K distributed, 13–15K contributors (Arena About). Divide it out and you get ≈$3.77 per successful break and ≈$0.12 per attempt [UNVERIFIED — this is my arithmetic on company-reported aggregates, and it excludes platform, staff and grading cost]. Buyers are named on the sponsor list: OpenAI, Anthropic, Google DeepMind, Meta, Amazon and UK AISI. The observable transaction is a sponsored competition at $20k–$170k+ per event, which is a fixed-price corpus commission wearing a tournament costume.
Defensibility: low. At under four dollars a unit, the corpus is not the moat. The crowd is portable and the tooling replicable; Gray Swan's actual moat is grading discipline — "automated graders and human judges, including independent evaluators, to prevent reward hacking" (Latent Space) — plus community gravity. A cyber-specialist version is more defensible than a general-jailbreak version only because the contributor pool is credentialed and small. See The bounty platforms and Paying the crowd.
Expert reasoning traces on vulnerability triage and exploit development. One annotated triage or exploitation trace. This is the artefact with the strongest demonstrated demand anywhere in the dossier: Anthropic's six external firms reviewed 5,008 findings, and the published agreement statistics — 85.2% exact at programme scale, 89% on the 198-report sample — are themselves the product, because they are what lets a lab claim its model's severity judgements are calibrated (Anthropic CVD). Price anchors: $80–90/hr, $1,750–2,150 per completed pentest task, and FrontierMath's $300–1,000 per authored problem as a per-artefact comparator. Defensibility: high — credentialed reviewers with reproduction capability are the scarce input.
Detection rules and detection-engineering rationale. One Sigma, YARA, Sentinel or Splunk rule, plus the why. SOC Prime's Threat Bounty is the only real per-artefact detection-rule market that exists, publishing 81 community rules in October 2024 alone, with monetisation "exclusively depend[ing] on how useful and actionable content is considered by organizations leveraging the SOC Prime Platform" (SOC Prime). Its per-rule payouts are not published and I could not establish them [COULD NOT ESTABLISH] — a notable hole, given it is the closest thing to a comparable. By analogy to the $200–2,000/task RL band, a rule-plus-rationale with a synthetic-telemetry verifier plausibly sits at the low end, $200–800 [UNVERIFIED]. Defensibility: medium — finished rules commoditise fast (SigmaHQ gives away 3,000+ under a licence that permits resale), but the rationale, the rejected drafts and the false-positive profile do not, and that is exactly what a model needs. See Defensive security.
Malware reverse-engineering walkthroughs. One binary plus a step-by-step analysis trace. Crystal Peak's rust_vm challenge for AISI is the proof this is contractable. Verifier quality is medium: you can check recovered structure against ground truth if you authored the binary — which pushes toward synthesised samples over real malware, and neatly sidesteps the licensing questions hanging over VirusTotal and MalwareBazaar redistribution. Pricing [UNVERIFIED]: reverse engineering is the highest-skill, lowest-supply labour in the stack, so the top of the $75–200+/hr expert band and $2,000+ per completed artefact. Defensibility: high.
Threat-intel analysis with attribution reasoning. One assessed intrusion set with the reasoning chain and the confidence language. Verifier quality: none — attribution is a judgement. That makes it a pure service, and it is also the weakest model capability found anywhere: AthenaBench puts GPT-5 at 39.0% on threat actor attribution against 92.0% on knowledge questions (arXiv 2511.01144). Defensibility is high on the labour and zero on the artefact — a written trace is trivially copyable once delivered, so it must be sold with hold-out terms or as a retained panel. No transaction price for AI-training-oriented threat-intel data exists publicly [COULD NOT ESTABLISH].
Held-out private evaluation sets, and evaluation-as-a-service. The highest-margin shape available, and the best-evidenced. Irregular runs it: a private challenge library on their own infrastructure, connecting into the lab's model or scaffold, delivering "continuous benchmarking, as well as qualitative and quantitative assessments," with the evaluation set "private to avoid contamination" (Irregular). Vals AI runs the enterprise version, which Sacra describes as subscription plus usage-based evaluation volume plus human review workflow fees (Sacra). Defensibility: very high, for a structural reason that is rare in data businesses — the customer cannot take it in-house without destroying the property that makes it valuable.
Red-team-as-a-service. ARIA's is the cleanest public structure: one partner at ~£2–3m over ~15 months, executing "a thorough pen-testing effort of every Blue Team's system" within each 8-week sprint, delivering "a detailed report covering attack attempts, outcomes, and reproducible artifacts" (ARIA). That works out at roughly £160k–£200k a month for a team — an engagement rate, not an artefact price. Gray Swan's private arenas are the commercial equivalent: ~20 handpicked Arena veterans against a customer's proprietary agent, with "a pricing and incentive structure upfront" that "differs for each arena." No published day rate for AI-lab-facing red-team work exists [COULD NOT ESTABLISH]. See Adversarial evals and red-team crowds and Government buyers.
Trained reward models and graders. A grader that replaces human review. This is the endgame artefact and the one that converts a consultancy into a data business: those 85.2%/89% severity-agreement statistics are, functionally, a training and validation set for a severity grader, and Anthropic's 26,153-candidate funnel shows exactly the bottleneck it would relieve. Defensibility: high if the underlying trace corpus is exclusive, near-zero if not — a grader is a small artefact and easy to distil once it is out. No public price for a domain grader exists [COULD NOT ESTABLISH].
Contract shapes, and the five terms that decide the outcome
| Shape | Real instance | Terms |
|---|---|---|
| Fixed-price service with milestone payments | ARIA Safeguarded AI red team | ~£2–3m / 15 months, 8-week sprints, acceptance criteria, explicitly "commercial services terms (not R&D grants)" |
| Quarterly retainer | General RL-environment market | "Six to seven figures per quarter"; engagements at least a quarter long |
| Per-task / per-artefact | Mercor pentest tasks; Prime Intellect bounties; FrontierMath | $1,750–2,150/task; $1,000–5,000+/environment; $300–1,000/problem |
| Hourly, expert-graded | Mercor, Handshake cyber roles | $54–$111/hr |
| Commissioned benchmark, licence + hold-out | OpenAI ↔ Epoch, FrontierMath | Commissioner owns the questions; 300 with solutions, 50 statements-only hold-out; producer keeps the right to evaluate and publish |
| Sponsored competition as corpus commission | Gray Swan Arena | $20k–$170k+ prize pool per event, sponsor named |
| Hosted platform subscription | Irregular; Vals AI | Subscription + usage volume + human review fees; demo-led, negotiated, unpublished |
| Resource-in-kind, no cash | METR | "METR has not accepted funding from AI companies"; labs "provided access and tokens" |
Exclusivity windows. The 4–5x premium is the only price, and no published example of a time-boxed exclusivity window in an AI-data contract exists [COULD NOT ESTABLISH]. Proposing one — 12–24 months, then non-exclusive resale — is the obvious pricing lever, because it converts a 4–5x one-off into 1x exclusive plus a long non-exclusive tail. The mechanics are in The oracle decides everything.
IP assignment versus licence. FrontierMath is assignment-in-substance with a licence-back. ARIA is the opposite: implementations and tools stay with the builder, though ARIA "may reassign commercialization rights if no deployment occurs within 12 months." Default to licence. Assignment destroys the corpus asset, and the corpus asset is the entire valuation case — a distinction GMV is not revenue makes in revenue terms and this makes in balance-sheet terms.
Hold-out obligations. FrontierMath's 50-problem statements-only hold-out is the model, and it matters more in cyber, not less: an environment whose flags sit in the buyer's training set measures nothing.
Non-compete with rival labs. Not observable in any public contract text [COULD NOT ESTABLISH]. The proxy evidence points strongly against broad vendor-level non-competes: Gray Swan lists OpenAI, Anthropic, Google DeepMind, Meta, Amazon and AISI simultaneously, and Irregular serves OpenAI, Google and Anthropic at once. Exclusivity is scoped to the artefact, not the vendor.
Audit rights and security terms. Slator's survey of how labs choose data partners puts provenance and licensing clarity first and security protocols third — "controlled environments with strict access management for sensitive model snapshots" (Slator). SOC 2 Type II is a gate, not a differentiator: only 10 of 38 tracked vendors hold any SOC 2, and the Mercor breach — reportedly four terabytes including contractor ID documents and biometrics, March 2026 — is why labs now audit.
The artefacts that price well here are the ones with no oracle. The ones with a perfect oracle are being given away by the hundred thousand.