Miju Labs

The security dossier

AI red-teaming

The only sub-market with a public procurement record — and that record shows roughly 100 hours from one vendor and 16 from another per model release. Five-figure engagements, not seven-figure contracts, in the most crowded and fastest-saturating corner of security.

crowdedhigh confidence8 minupdated 2026-08-30

This is the only sub-market in the dossier where you can read the invoice, and reading it is chastening.

The Claude Opus 5 System Card names its external testers and counts their engagement — the single best procurement disclosure in AI safety (system card):

PartyEngagement
Trajectory Labs, PBC"roughly 100 hours red-teaming our safeguards"
10a Labs"around 16 hours testing"
Gray Swan"150 attempts per task"
UK AI Security Instituteopen-ended cyber-range testing, 100M token budget per attempt
IrregularCyScenarioBench, "measuring multi-stage cyber operations under realistic constraints"

One hundred hours and sixteen hours are small numbers. At even $250/hr those are engagements worth roughly $25,000 and $4,000. If they are representative of what a lab buys per model release from a single external safeguards red-team vendor, the per-lab-per-release revenue from hour-billed red-teaming is five figures, not seven. Everyone in this market sizes it off the logos. The logos are worth less than the hours suggest.

The money that is visibly larger sits in the range and benchmark relationships — UK AISI and Irregular — not in the safeguards hours. That is the strategic read: the sellable thing is the harness, not the attempt. See The oracle decides everything.

What the artefact is

An attack — a jailbreak, an indirect prompt injection, an agentic misuse chain — plus its trajectory, the target's response, and a severity grading.

The corpus that gets built from those is worth more than the individual attacks, and one company has already built it. Gray Swan's Arena reports 4M+ attack attempts, 130K+ successful breaks, 13K+ community members and $490K+ in rewards distributed, with 100+ participants placed in paid roles (Arena About). The company's own marketing copy elsewhere says "over three million" attempts and "15,000+ adversarial researchers" — date drift, not a discrepancy.

Do the division. $3.77 per verified break. $0.12 per attempt. $32.67 of lifetime reward per member. And the terms are total: "By submitting an entry, entrants grant Gray Swan AI and its partners an irrevocable, worldwide license to use and share the submission for any purpose." Gray Swan acquired perpetual sublicensable rights to 130,000 verified attack trajectories for under half a million dollars. That is the price this artefact clears at, and it is the number a new entrant has to beat. Paying the crowd and Gray Swan, in full have the rest.

Structurally, note that the target model belongs to the buyer. That makes this a service relationship by construction — you cannot hold the asset.

Is there a verifier

Mixed, and degrading. Split it by what is being attacked.

Where the target is a system, there is a flag. AIRTBench works because its 70 challenges come from Dreadnode's Crucible platform and are black-box CTFs — the model writes Python to compromise an AI system and either recovers the flag or does not (arXiv 2506.14682). AgentHarm works because scoring requires jailbroken agents to maintain their capabilities after the attack and complete a multi-step task — a capability check, not a vibes check, across 110 malicious agent tasks in 11 harm categories (arXiv 2410.09024).

Where the target is a safety policy, the verifier is a judge model or a human panel. StealthBench's method — a three-model LLM judge panel with majority-vote aggregation — is now standard, and it inherits every weakness of LLM judging.

[UNVERIFIED — judgement] As models improve, judge-based grading of jailbreaks becomes circular: the thing grading the attack is the class of thing being attacked. The durable artefacts in this sub-market are the flag-bearing ones, which is why the AIRTBench/Crucible shape and the dockerised-scenario shape are the ones worth building. Everything graded by panel is a service with a shelf life. Compare Malware reverse engineering, where the answer is a fact about a binary and stays one.

Is anyone buying

Budget scores 4. Demonstrably yes, at good bands — but with the caveat above about size.

BuyerRoleBand
AnthropicRed Team Engineer, Safeguards$320,000 – $405,000
AnthropicLead, Frontier Red Team (Cyber)$485,000 – $755,000
OpenAISecurity Researcher, Agentic AI Threats (Preparedness)$293K – $405K
OpenAIModel Policy, Frontier Cyber Risk$266K – $335K
Scale AIStrategic Projects Lead, Red Teamnot disclosed

And the government side is bigger than the lab side, which nobody expects. Gray Swan billed UK AISI £702,730 in a single day — two invoices on 31 March 2025 under DSIT's "Safety – Cyber – Consultancy" cost code. That is an order of magnitude above the safeguards hour counts, from a buyer who is not a frontier lab. What the labs pay sets out the full transparency-data table; Government buyers sets out the procurement shape.

Mercor's AI Red-Teamer — Adversarial AI Testing listing (now closed) offered $54–111/hr for exactly this work — "jailbreaks, prompt injections, misuse cases, exploits", "annotate failures, classify vulnerabilities, and flag systemic risks", restricted to the US, UK and Canada (listing). The $54 floor is below every other security rate in this dossier. The dispersion is the widest of the eight because the bottom of this crowd genuinely is low-skill.

What the expert costs

Cost scores 4 — the second-cheapest pilot after Governance, risk and compliance. Nothing to license, since attacks are generated rather than sourced. No estate to run, unlike Incident response and forensics. No corpus to buy, unlike Malware reverse engineering. A cohort, a target endpoint, a scoring harness.

The labour is cheap at the bottom and unremarkable at the top. A Gray Swan Red Team Engineer is paid $110,000–$185,000, which at full utilisation is about $89/hr before overhead — almost exactly the Mercor senior contractor rate, and less than half what the elite offensive tier commands in Vulnerability research.

Getting to them

Reach scores 3, and the reason is uncomfortable: the two best channels belong to competitors.

Gray Swan's Arena (13,000–15,000 members, ~15,000 on Discord) is the densest concentration of this population anywhere, and it is the incumbent's balance sheet. Dreadnode's Crucible is the second qualified register, and AIRTBench is built on it. The labs run the third: Anthropic pays up to $35,000 per novel universal jailbreak and OpenAI up to $100,000 for exceptional findings, which is direct disintermediation of any vendor sitting in between.

The population combining adversarial security with deep LLM and agent knowledge is estimated at "the low thousands globally" [WEAK — blog citing hiring-leader conversations], corroborated by Gray Swan's roughly 0.7% crowd-to-professional conversion rate. Small, already assembled, already under contract. See Offensive supply.

Where the benchmarks sit

BenchmarkScaleFrontier score
AIRTBench70 black-box CTF challenges (Dreadnode Crucible)Claude 3.7 Sonnet 43/70 (61%), 46.9% overall; Gemini 2.5 Pro 39 (56%); GPT-4.5-Preview 34 (49%); best open-source Llama-4-17B 7 (10%). Prompt injection 49%; system exploitation and model inversion <26%
AgentHarm110 malicious agent tasks (440 augmented), 11 categoriesLeading models "surprisingly compliant" without jailbreaking; simple universal templates transfer
StealthBench11 hand-verified OPSEC incidents → 14 dockerised scenarios≤54% safe success rate

AIRTBench's most important number is not a score. It reports that models complete challenges "in minutes what typically takes humans hours or days — with efficiency advantages of over 5,000x on hard challenges."

The benchmark measures its own suppliers out of the market

A 5,000x efficiency advantage over human attackers, stated inside the benchmark that defines the discipline, is a direct measurement of human AI-red-teaming labour being economically displaced. Every other sub-market here has a capability gap that argues for buying more expert hours. This one has a published multiple arguing for buying fewer.

That is why defense scores 2 and room scores 1. Gray Swan, Irregular, Dreadnode, Trajectory Labs and 10a Labs are all in the named procurement record; the labs run their own bounties; and the crowd's output clears at $3.77 a break.

What would kill it

What would kill it

Automation from inside. The 5,000x figure is not a forecast. It is a measurement, published by the benchmark's own authors.

Direct bounties. A lab paying $35,000–$100,000 for a novel jailbreak has no structural need for a vendor between it and the researcher.

The hour counts stay small. If 100 hours per release is the steady state per vendor, five external firms split a market that is a rounding error against a single UK AISI invoice.

Where the record is thin

Proof scores 2 because Gray Swan is cited in 11 recent frontier model system cards by its own Series A release, and lists 17+ on its site. Displacing that reference position is not a one-paper exercise.

Nothing is public about what Trajectory Labs or 10a Labs were paid — the system card gives hours, not money, and the $250/hr conversion above is an assumption, not an observed rate. No AI red-teaming contract value is published anywhere except the UK AISI transparency line. And Gray Swan's own revenue mix is undisclosed and cannot be derived; the most that can be said is that it runs two quota-carrying segments with six-figure-plus deals across roughly 50 people. See Gray Swan, in full and the ranked read in The security read.