The question is narrow and answerable: does publishing a cyber benchmark convert into revenue? The evidence says yes — and only when a commercial vehicle is already attached to catch it. Citations alone convert into influence, which is a different asset and does not pay salaries.
The positive cases
Vals AI is the cleanest. Sacra describes the mechanism without hedging: "Public benchmark results attract media coverage from outlets like the Wall Street Journal, Bloomberg, and Law.com, building trust with enterprise buyers and model labs" — and then the sentence that matters, "public benchmark investment is not just a marketing cost, it is the mechanism that makes private enterprise evals credible enough to sell" (Sacra). The outcome: $40M Series A at a $400M valuation with revenue up 8x, a16z-led, August 2026 (Pulse 2.0).
Note the sequence carefully, because it is the template. Public benchmark → media citation → inbound enterprise → private evaluation contract on the customer's own data. The public benchmark is never itself the revenue line. It is the credential.
Irregular is the cyber instance of the same pattern. Their research index carries a continuous stream of model-by-model offensive-security assessments — GPT-5.4, GPT-5.5, GPT-5.6, GLM-5.2, Kimi K3, Muse Spark, Claude Sonnet 4.5, Claude 3.7 Sonnet (irregular.com/research). They are cited in Anthropic's and OpenAI's model documentation, and Anthropic named them as the collaborating partner on the cyber-evaluation incident review (Anthropic). The outcome: $80M at a $450M valuation, with "millions of dollars in annual revenue" at the time of the raise (SecurityWeek).
Time from first public model evaluation (Claude 3.7 Sonnet, early 2025) to Series B (September 2025) was roughly six to nine months [UNVERIFIED — inferred from publication dates on their research index, not stated by the company]. That is the fastest conversion anyone in this market has demonstrated, and it should be read as a ceiling rather than a plan.
Gray Swan converted differently, and the difference is instructive. The Arena is a marketing engine, a corpus generator and a hiring funnel simultaneously; the revenue products are Shade and Cygnal. It appears in 11 frontier model system cards, and it has placed 100+ Arena participants into paid red-teaming roles (Gray Swan). $40M raised. The benchmark here is not the credential for a private eval — it is the recruitment channel for the labour that makes the product.
The control case
METR is the most-cited evaluation organisation in the field and it takes no lab money. Its own about page: METR "has not accepted funding from AI companies." Labs provide "access and tokens" only. The one contract line is "a small part of our income… from a technical assistance contract with the European AI Office" (METR).
Citation volume is not a business model. It converts to influence — real, durable, arguably more valuable than revenue for a nonprofit — but it converts to revenue only through a commercial vehicle that exists to catch it. METR deliberately does not have one, and the result is an organisation with maximum standing and near-zero commercial capture.
The corollary for anyone building here: decide the vehicle before publishing the benchmark, not after. The benchmark is easy to publish and the credibility it buys decays if there is nothing to spend it on.
Epoch AI sits between, and paid for it. FrontierMath was the revenue — OpenAI commissioned it directly. But Epoch could not disclose the funding before the o3 announcement because of the contract, and contributing mathematicians were not told OpenAI had exclusive access (Epoch; TechCrunch). Direct commissioning converts fastest and damages credibility fastest. For a cyber specialist whose entire asset is trusted judgement, that is a worse trade than it was for Epoch — the product being sold is the belief that the grader is not captured.
Cybench is the academic default outcome. It came out of Stanford, became infrastructure — packaged inside UK AISI's Inspect Cyber (AISI) and re-run commercially by Cotool (Cotool) — and its authors captured the citations while third parties captured the revenue. That is what happens to a benchmark released without a vehicle attached.
What it costs to build the credential
The published inputs, assembled:
| Build | Effort actually disclosed |
|---|---|
| Cybench | 40 professional CTF tasks sourced from four existing competitions (HackTheBox Cyber Apocalypse 2024, SekaiCTF 2022–23, Glacier, HKCert), first-solve times 2 minutes to 24h54m, decomposed into subtasks, containerised (arXiv 2408.08926) |
| AIxCC | 48 challenge projects / 63 vulnerabilities from 24 OSS repos; ground-truth validation consumed 8,906 CPU-hours of reference fuzzing plus two person-weeks of manual annotation by two independent reviewers plus $341 of Claude API; each finalist team got $85k cloud + $50k LLM credits (arXiv 2602.07666) |
| AISI's ranges | ~20 and ~15 human-expert solve hours, built by three separate external firms on bare-metal nested virtualisation (AISI) |
| CTI-HAL | 2 annotators, 81 reports, 8 weeks — ~5 reports per annotator-week (arXiv 2504.05866) |
The Cybench paper publishes no cost, effort or funding figures at all [COULD NOT ESTABLISH] — which is unfortunate, because it is the most widely reused cyber benchmark in existence and therefore the single most useful cost comparable there could be.
What the paper does reveal is the shortcut, and it is the important part: sourcing tasks from four existing CTF competitions collapses the authoring cost to near zero. The price paid is contamination risk — a task that has been public and solved is a task that may already sit in the training set. That trade is the whole design decision behind cheap benchmarks, and it is why cheap benchmarks stop measuring anything within a year.
From those inputs, a construction [UNVERIFIED — my synthesis, not a sourced figure]:
| Tier | Realistic cost | What the money buys |
|---|---|---|
| CTF-derived, 40 tasks, flag oracle (Cybench shape) | $50k–$150k | Mostly engineering and harness; authoring largely avoided by sourcing existing CTFs |
| Purpose-built multi-host range, one scenario (AISI TLO shape) | $100k–$400k | 200–1,000 expert-hours at $80–90/hr, plus bare-metal infrastructure, reset tooling, red-herring design, two-plus QA cycles |
| Open-ended vulnerability research on real targets (FrontierCyber shape) | $500k–$2m+ per year, ongoing | Continuous re-authoring as models saturate, a physical device estate, a standing expert review panel |
Two external anchors sit comfortably inside that frame, which is the main reason to have any confidence in it: ARIA's £2–3m over ~15 months for one red-team partner, and Epoch's $300k for a complex product replica. The build-cost mechanics are in The oracle decides everything and the sale-side mechanics in What you can actually sell.
The binding constraint is not the money. AISI's own range designer states it: "people with these skills are busy defending actual systems" (FAR.AI). Compare Defensive supply and Offensive supply on how thin that pool actually is.
Contamination is what makes the model work
The reason this is not simply "spend money on marketing" is that in cyber, publication actively destroys the thing published. A benchmark whose tasks are public is a benchmark that may be in the next training run — the effect CrowdStrike names directly when it criticises public benchmarks for saturation and contamination.
That destruction is the business model, not a bug in it. Publish the benchmark; hold back the private version; sell the private version.
Irregular does exactly this. Their evaluation set "remains private to avoid contamination" and their platform advertises "proprietary in-house cyber challenges (avoiding training data contamination)" (CyScenarioBench). Their published model assessments are the shop window; the held-out library on their own infrastructure is the shop. Vals AI runs the identical structure in enterprise evaluation. FrontierMath encodes it contractually: 300 problems with solutions delivered, 50 statements-only held out.
The public artefact and the sellable artefact must be different objects with a deliberate gap between them — and the gap is what a buyer pays the exclusivity premium for, which The oracle decides everything puts at 4–5x in the general RL market with no cyber-specific figure established.
This also creates the rarest moat in the data business: a customer who takes the private evaluation set in-house destroys the contamination protection that made it worth buying. They cannot leave with it. That is worth more than any of the citation counts on this page.
One structural tailwind is worth naming. DARPA deliberately converted $9.9M of public money into permanently free, open-source cyber reasoning systems and an open challenge corpus at AIxCC (DARPA). Anyone selling vulnerability-discovery tooling now competes with free. Anyone selling novel, unpublished, expert-authored evaluation content benefits, because the open corpus is now public, contaminated for benchmark purposes, and useless for measuring a model that may have trained on it.
The open-sourcing of AIxCC increased the value of private, held-out, expert-generated evaluation data. So did Anthropic saying out loud that it can no longer track capability progression with the benchmarks it has. Both are demand statements from the same direction, and the measurement picture in The measurement gap says where the remaining unsaturated ground is: not offensive, where the curve has already run, but the defensive half nobody has properly measured. Note also what publishing does not buy — see GMV is not revenue before treating benchmark-driven inbound as revenue quality.