Code is the only domain in this atlas where the buyer's budget is proven, enormous and already well supplied. That is the whole problem.
SWE-bench Verified is effectively saturated. The August 2026 top five — Claude Opus 5 at 96%, Claude Mythos 5 at 95.5%, Claude Fable 5 at 95%, Claude Opus 4.8 at 88.6% and Claude Opus 4.7 (Adaptive) at 87.6% — sit clustered within 1.0 points at the top (BenchLM). A benchmark that cannot separate the leading models has stopped being a purchase argument.
Meanwhile the benchmarks for judging code are months old and barely populated: the Code Review Agent Benchmark (arXiv 2603.23448) and AACR-Bench for repository-level automatic code review (arXiv 2601.19494). That is the entire public corpus for the sub-skill.
The asymmetry is the thesis. Writing code is solved-ish and served by roughly twenty funded vendors. Judging code — architecture critique, "this PR is technically correct and strategically wrong", legacy-system archaeology, root-cause narration on a real production incident — has near-zero benchmark coverage and no dedicated vendor. It is a real niche inside a crowded lane, which is a different thing from a real market.
What the data actually is
Four artefacts, and only the last two are hard to copy.
Review comments with rationale. A merged PR, the reviewer's comments, and — the part that does not exist in the wild — the reasoning behind the comment that never got typed because the reviewer assumed it was obvious.
Architecture critique. A design document plus a senior engineer's read on why the proposed approach will be fine for eighteen months and then expensive. The judgement is temporal, which is exactly what a model trained on merged diffs cannot see.
Legacy archaeology. Someone reconstructing why a fifteen-year-old subsystem is shaped the way it is, from commit history and nothing else. Narrated. This is the narrated-trajectory form and it is the most valuable item on the list.
Root-cause narration. A real production incident, walked through by the person who resolved it. Note that this straddles Defensive security — the incident-response narration product there has the same shape and a much emptier benchmark.
Proof scores 3: the benchmark vacancy is real, but so are twenty vendors with existing lab relationships who can occupy it faster than you can publish.
The harness sensitivity finding is worth building a product around on its own: the same Claude Opus 4.5 scores 80.9% on SWE-Bench Verified and 45.9% on SEAL (Kili survey). A 35-point swing on harness alone tells a buyer their headline number is not measuring what they think.
Is anyone buying
Yes, more than in any other niche here, and budget scores 5 without argument.
SWE-bench Pro, built and maintained by Scale AI — 1,865 tasks across 41 professional repos, averaging 107.4 changed lines across 4.1 files — still has headroom: GPT-5.4 (xHigh) reaches 59.1% on the public set, and on the commercial proprietary set the best score is under 48% (Scale, Morph). That headroom is where the coding-data money currently goes.
SemiAnalysis reports OpenAI paying roughly $20,000 per website for agent-training environments [WEAK — trade newsletter] (SemiAnalysis). And the labs are visibly moving past general code into named sub-verticals: Anthropic is hiring Research Engineer, Cybersecurity RL and Research Engineer, Chip Design RL (Anthropic) — RL role titles that name a domain, which is the shape of the next budget line.
Those RL engineer roles hire engineers to build training systems, not practitioners to produce review data with a published band. The named data-producing reqs with pay bands in this atlas are cyber (OpenAI's Red Team Specialist — Cyber, $198K–$320K, "constructing datasets") and health ($310K–$460K). Code has enormous budget and no posting that says "we pay senior engineers to critique PRs" — the spend flows through vendors, which means your buyer is a procurement process you cannot see.
Getting the experts
Reach scores 5, and this is the domain's genuine structural advantage over every other niche in the set.
GitHub commit and review history is a public, verifiable credential for exactly this skill. You can rank candidates by real review comments on real merged PRs — not by a portfolio, not by a licence, not by a self-report. No other domain permits that. Law gives you a bar register that proves admission but says nothing about drafting quality; Design and UI/UX gives you a portfolio that shows output but not judgement. GitHub shows the judgement itself, at scale, retrospectively, for free. See Who is actually on the other end on why a credential that cannot be faked cheaply is worth more than a licence.
Other channels: maintainer lists of major OSS projects, the Rust, kernel and CPython mailing lists, Hacker News, r/ExperiencedDevs, Lobsters, and conference speaker lists.
The pool is 1,717,800 US software developers at a median $135,980/yr (~$65.28/hr), growing 10% (BLS); GitHub claims roughly 180M developers globally [WEAK — trade summary of Octoverse 2025] (GitHub blog).
What it costs to run
Cost scores 4. Turing pays software engineers $60–120/hr and Mercor's expert band runs $50–200/hr [WEAK — aggregator] (theairankings). Against a $65.28/hr BLS median, the arbitrage is thin — roughly 1x to 1.8x the day job, versus 2–3.5x in some other niches — because this is the one profession that has already been bid up by every data vendor in the market.
Capital requirements are otherwise zero: open-source repositories, public issue trackers, a CI runner.
Who is already there
Almost everyone. Of ~37 vendors in the RL-environment directory, the plurality are coding-focused: Mechanize, AfterQuery, Bespoke Labs, Huzzle Labs, Datacurve, Proximal, BenchFlow, Refresh, Vmax, Habitat, Scale, Surge, Mercor, Turing (rl-list).
Datacurve is the closest pure vertical play — a $15M Series A led by Chemistry after a $2.7M seed from Balaji Srinivasan, with angels from DeepMind, Vercel, Anthropic and OpenAI. It pays engineers through a bounty-hunter model and has distributed over $1M in bounties; co-founder Serena Ge: "We treat this as a consumer product, not a data labeling operation" (TechCrunch). Mechanize raised $9.1M (Mechanize). AfterQuery raised a $30M Series A at $300M post past a $100M annualised revenue run-rate — gross, on ~100,000 verified professionals — and claims every US frontier lab as a customer (Sacra).
Room scores 2. An entrenched specialist caps room at 2 however good the rest looks, and here there are twenty of them.
What would kill it
The sub-lane is a feature, not a company. Datacurve, AfterQuery or Scale can add review-and-critique tasks in a sprint using supply they already have. The review benchmark you publish becomes their roadmap item.
Employer IP assignment. Employment agreements assign IP in real work code, so proprietary review threads cannot leave the employer. You are confined to open source — which is fine, and also identical to what every competitor is confined to.
Defense is 2 and that is the honest number. 1.7M developers, no licence, no register, and a pool every vendor in the market is already courting. Anyone can run your recruiting campaign next month.
Saturation moves. If code review benchmarks follow SWE-bench Verified from 22% to 96% in three years, the window is short.
The first ninety days here
If you enter at all, enter narrow: one language, one class of judgement, one benchmark. Mine GitHub for the 200 reviewers with the highest-signal comment histories in that language, approach them individually, and pay per artefact rather than per hour — Datacurve's bounty model works in this profession for the same reason it works in Offensive security: engineers accept per-artefact payment from strangers.
Publish an architecture-critique benchmark where the ground truth is what happened to the codebase eighteen months later. That is defensible because it needs time to construct and cannot be synthesised, and it sidesteps the AACR-Bench comparison entirely. See The first ninety days and The specialist wedge — code sits on the side of the trade where the money is proven and the benchmark is captured.
Where the record is thin
The ~20 coding vendors figure is a count from directories that are SEO properties rather than registries, and their funding data is unreliable.
The $20,000-per-website environment price is a single trade-newsletter claim. Turing's and Mercor's rate bands come from a self-reported aggregator and should be treated as directional. And no one has published what a review-specific dataset sells for, because no one has sold one — the entire pricing case here is extrapolated from general coding-data rates.