Miju Labs

All niches

Offensive security

The strongest demand evidence of any niche in this atlas — named in system cards, with lab reqs carrying pay bands — which is exactly why two funded specialists already own the network and the benchmark.

crowdedhigh confidence7 minupdated 2026-08-30
Who the expert is
Pentesters, exploit developers, CTF players, bug-bounty hunters
What they earn by day
$62.11/hr median (BLS, information security analysts); pen testers average $43/hr
What data work pays
Hourly, and higher than this page originally assumed — Mercor posts $200–250/hr for offensive security and vulnerability research; Gray Swan's tournament acquires a break for ~$3.77
Size of the pool
192,900 US information security analysts; Gray Swan's Arena alone holds 15,000+ researchers
How you reach them
HackerOne and Bugcrowd leaderboards, CTFtime rankings, DEF CON villages, OSCP register, r/netsec
Benchmark position
Vendor-captured — Gray Swan cited in 11 system cards; Irregular's SOLVE and CyScenarioBench appear in them too
Who is already there
Gray Swan, Irregular, Dreadnode, Trajectory Labs, 10a Labs, Palisade Research
Read
Ranked first on evidence and last on availability. The benchmark layer has already been captured by vendors, which makes the usual entry weapon useless here.
Speed to proof
2
Budget now
5
Defensibility
3
Cheap to start
3
Reachability
5
Room to win
2
This page has been superseded — read The security read first

The deep dossier reframed the whole domain: offensive versus defensive is the wrong axis, and whether the artefact has an automated oracle is the right one — memory-corruption environments are already a commodity Bugcrowd ships by the hundred thousand, while validating logic bugs has no oracle at all and Anthropic pays six outside firms to do it by hand. Most of the scoring below survives; the cost score does not, and is argued down at the end of this page. The replacement scoring is at The security read.

Offensive security is the niche with the best evidence and the worst timing, and the two facts are the same fact.

Every argument a founder makes about why a domain will work is already demonstrated here. Labs buy: OpenAI's Red Team Specialist — Cyber posts $198K–$320K plus equity, with responsibilities that explicitly include "constructing datasets" and designing evaluation frameworks for model cyber capabilities (OpenAI). Anthropic is hiring Lead, Frontier Red Team (Cyber), Research Scientist, Frontier Red Team (Cyber) and — note the title — Research Engineer, Cybersecurity RL, which is a training-data role, not a safety one (Anthropic).

Vendors are named in the artefacts that decide procurement. Claude Opus 5's system card names, for cyber alone: the UK AI Security Institute, Irregular (CyScenarioBench), UC Berkeley / Max Planck / UC Santa Barbara / Arizona State (ExploitGym), Mozilla, Trajectory Labs PBC, 10a Labs and Gray Swan. Disclosed cyber scores include ExploitBench at 10.14 mean flags and 99 full ACE exploits, OSS-Fuzz 79.4% non-zero, Firefox 147 at 52.4% full exploits (131/250) and CyScenarioBench 33.7% completion (Claude Opus 5 System Card).

That is the complete demand proof every other page in this atlas is trying to assemble. It is also the reason two funded companies are standing in the doorway.

What the data actually is

Attack trajectories, and the profession is already accustomed to selling them.

This is the only domain here where practitioners are used to being paid per artefact by a stranger on the internet. Bug bounty and CTF prize norms mean you can pay $1,000 for a good exploit trajectory rather than $43/hr for someone's time — fundamentally better unit economics than any hourly domain, and Gray Swan AI has already proven it at 15,000 people. Its Ultimate Jailbreaking Championship offered $40,000 total: $1,000 per model for the first jailbreak of the first 20 models, $2,000 each for the final five, plus $1,000 each to the top 10 participants, with top performers considered for employment (LessWrong announcement).

The artefacts themselves: full attack trajectories against synthetic targets, jailbreak transcripts with the reasoning preserved, exploit development narrated step by step, and human baselines — Anthropic entered Claude into seven cybersecurity competitions in 2025, from picoCTF to CCDC, placing in the top 25% in many, explicitly to obtain human baselines against professional researchers, thanking Palisade Research, WR CCDC, Plaid Parliament of Pwning, DEF CON Qualifiers and Airbnb (Anthropic).

Is anyone buying

Yes, harder than anywhere else in this atlas. Budget scores 5 and it is the only uncontested 5 in the set.

Beyond the reqs and the system card above, the vendors' own funding tells the story. Gray Swan AI raised a $40M Series A co-led by Wing Venture Capital and Madrona, with Obvious Ventures, Snowflake Ventures, Hudson River Trading, Samsung Next and Magarac; it is cited in 11 recent frontier model system cards from Anthropic, OpenAI and Meta, serves 20+ customers, and its Arena community of 15,000+ researchers has generated over one million high-quality real-world attack trajectories (National Law Review press release). A $200M valuation is reported by one secondary outlet [WEAK] (Silicon Report).

Irregular — formerly Pattern Labs, founded 2023 in Tel Aviv by Dan Lahav and Omer Nevo, roughly 35 employees — raised $80M at a $450M valuation from Sequoia and Redpoint in September 2025 and runs cyber evaluations for OpenAI, Anthropic, Meta and Google DeepMind (TechCrunch, BERI). Three competing labs buy "independent" assurance from the same 35-person firm — a One customer is a binary event fact worth sitting with.

Getting the experts

Reach scores 5 and the channels are unusually honest: they rank people publicly by the exact skill you are buying.

HackerOne and Bugcrowd leaderboards are ranked and self-selecting. CTFtime team rankings order the competitive population globally. Then DEF CON villages and Black Hat Arsenal, HackTheBox and TryHackMe ladders, the OffSec OSCP/OSCE certification register, university CTF teams like PPP, r/netsec at roughly 460,000 members [WEAK — third-party tracker] (Reddgrow), the Bug Bounty Forum and Hacker101 Discords — and Gray Swan's own Arena, which is simultaneously your competitor and an index of the pool.

BLS counts 192,900 US information security analysts at a median $62.11/hr, growing 21% (BLS). DEF CON attendance is last reliably reported at ~30,000 for DEF CON 27 in 2019 [WEAK — stale] (Wikipedia).

What it costs to run

Cost scored 4 here originally and is now 3 — the argument is at the end of this page. US penetration testers average $88,575/yr ($43/hr), with the 90th percentile at $100,587 and expert-level 8+ years exceeding $172,000 (Salary.com) — but that salary average is not what this work clears at.

But the hourly rate is the wrong frame, which is the good news. Paying per artefact means you pay only for the exploits that land, which converts your largest variable cost into something that scales with output quality rather than with hours logged. Capital requirements are low: synthetic targets, sandboxes, and a leaderboard. Gray Swan's championship shows the whole cost structure — $40,000 bought a competitive event that produced usable trajectories and a recruiting funnel simultaneously.

Who is already there

Room scores 2 because an entrenched specialist caps it there, and there are six.

Beyond Gray Swan and Irregular: Dreadnode raised $14M for security-agent infrastructure (SecurityWeek). Trajectory Labs, PBC runs contract AI red-teamer roles and states its work appears in Anthropic and Meta materials (Trajectory Labs). 10a Labs is hiring AI Red Teamer and AI Red Teamer (Cyber) and running a fellowship (Greenhouse). Palisade Research supplied HackTheBox data to Anthropic (Anthropic). See Adversarial evals and red-team crowds for the labour market underneath all of them.

The benchmark has already been captured

Proof scores 2, and this is why. The public benchmarks are open but not decision-relevant: NYU CTF Bench, 200 CSAW challenges, has Claude 4.5 Opus at 59.0% (118/200) and Gemini 3 Pro at 52.0%, up from 22% for Claude 3.5 Sonnet in prior work (arXiv 2604.17159) — genuinely unsaturated. Cybench, CyberSecEval and CyberMetric exist publicly.

The suites that actually move procurement — ExploitBench, ExploitGym, CyScenarioBench, SOLVE — are private or vendor-owned. A new public benchmark is a much weaker weapon here than in Law or Accounting, audit and tax, because the buyers have already standardised on someone else's. That is the exact trade The specialist wedge describes, seen from the losing side: the domains with proven budgets have captured benchmark layers, and the domains with empty benchmarks have unproven budgets.

What would kill it

What would kill it

Nothing kills the market; the market kills you. Gray Swan has 15,000 researchers and a million trajectories, Irregular has $80M and four lab customers, and both are named in the system cards your prospects read. Room is 2 for good reason.

Dual-use optics. Publishing exploit-development training data invites the "you are arming the models" objection, and Anthropic's own Opus 5 card already shows 52.4% full-exploit rates on Firefox 147. The safety story is a genuine commercial risk, not a talking point.

Clearances and NDAs. Real engagement reports are covered by client NDAs — hence synthetic targets like ExploitGym and CyScenarioBench. Security clearances exclude some of the best people from anything with a foreign-lab customer.

Concentration. If three labs buy assurance from one 35-person firm, a regulator's view on evaluator independence could reshape the buying process faster than any competitor could.

Defense scores 3: the supply stays scarce and the attack surface refreshes with every model release, but the pool is publicly ranked and anyone can run your recruiting campaign — Gray Swan's own Arena proves how quickly 15,000 people can be assembled by someone with a leaderboard.

The first ninety days here

If you enter, do not enter against the Arena. The uncontested ground is the seam with Defensive security — incident response, detection evasion narrated from the attacker's side — where the benchmark is empty and no vendor is named in any system card.

Otherwise: run one competition, not a recruiting funnel. $40,000 in prizes on CTFtime and HackerOne buys trajectories, a leaderboard and a hiring pipeline at once, and it is the only entry motion this profession responds to. Publish the results with the reasoning traces attached, which is the one thing Gray Swan's arena data does not make public. See The first ninety days.

One score the deep dossier moved, and one argument it kills

cost: 4 → 3. This page priced the labour off a $43/hr pentester average and then argued that paying per artefact converts the largest variable cost into something that scales with output quality. Both halves are wrong.

The labour does not clear at $43/hr. Mercor posts $200–250/hr for offensive security and vulnerability research, seeking top-ranked CTF team members, CVE discoverers and publicly recognised bug bounty researchers (Mercor); $100–150/hr for harm labelling; $85–95/hr for blue-team incident reasoning. The band is $54–250/hr, not a salary average — see The generalists in cyber. On top of that the deep look found costs this page did not count: SOC 2 Type II or ISO 27001 as a gate rather than a differentiator, a hold-out set that is inventory you build and never sell, and reproduction infrastructure with an incident-response posture.

And the per-artefact argument does not survive contact with the reference price. Gray Swan AI has distributed $490,000+ across 130,000+ successful breaks — roughly $3.77 a unit, with an irrevocable worldwide sublicensable licence attached (Arena About). You cannot beat that on piece-rate, and you should not try: a bounty crowd optimises for rare findings, while a lab training an agent needs coverage, reproducibility, consistent format and clean provenance. That is a payroll-shaped product, and it is the expensive one. Paying the crowd works it through.

The other five scores stand. proof: 2 in particular is confirmed rather than softened — the measurement layer is not merely captured, it is exhausted, with Anthropic stating Opus 4.6 "has saturated all of our current cyber evaluations" and retiring CyberGym by Opus 5 (The measurement gap).

Where the record is thin

Gray Swan's $200M valuation is single-sourced. DEF CON's attendance figure is seven years stale. The r/netsec membership number comes from a third-party tracker.

Nobody has published what a lab pays per attack trajectory, so the bounty economics above are inferred from one championship's prize structure. And no system card discloses purchased expert data as a line item — the vendor names appear as evaluation partners, which is a weaker claim than a contract.