Build, but in the specialties where the credential is not a medical degree, and sell the thing FDA wrote down rather than the thing the labs are buying today. The defensible business in health is blinded, adjudicated expert panels with measured disagreement, produced without ever touching a patient record — argued at Characterised disagreement. The four niches worth entering are Medical coding, Nursing, Dermatology and ophthalmology and Clinical reasoning, in that order of quality and the reverse order of revenue. This page is why the other nine are not, and why the screen was wrong about the thing it was most confident about.
What the deep look changed
The shallow screen scored clinical medicine well on evidence — HealthBench and MedHELM are open, the gaps are real — and killed it on cost, on the reasonable ground that physicians are the most expensive supply in the shortlist. That objection is correct. It applies to one specialty out of thirteen.
Radiology is the case the screen had in mind, and it is worse than the screen knew. Radiologists average $571,000 (Medscape via Becker's) — $286–317/hour of opportunity cost [INFERENCE]. The top observed AI-data rate for a radiologist is $300/hr on Handshake AI [WEAK] (aigigjobs), and teleradiology piecework already clears $250–400/hr while actively reading [WEAK] (natoe.ai). There is no arbitrage; there is a discount you are asking a radiologist to accept. See Radiology.
Now the other twelve. Certified coders cost $22–30/hour — BLS puts medical records specialists at a $24.59/hr median across 200,700 jobs (BLS) — against $80–110/hr observed at OpenTrain and Mercor (OpenTrain, Mercor). Registered nurses sit at $46.90/hr across 3,465,400 jobs (BLS) against $55–120/hr. Behavioural-health counsellors are at $28.53/hr against $80–150 — the widest multiple in the dossier.
The cost objection was never about medicine. It was about physicians. The arbitrage is real, large and repeatable wherever the credential is issued by a certification body rather than a medical school — and healthcare has more such credentials, better registered, than any domain in this atlas. Score cost against a coder rather than a radiologist and it goes from 2 to 5. The full ladder is at What a clinician hour costs.
The deep look also reversed the product. Models now beat physicians on HealthBench Professional, 59.0 against 43.7, and by four times on SDBench — MAI-DxO at 80% against 21 generalist physicians at 20% (OpenAI PDF, arXiv 2506.22405). You cannot sell ground truth to a buyer whose model already beats it. See The measurement landscape.
The finding that most threatens the business
OpenAI published its physician recruitment funnel in full, and there is no vendor anywhere in it.
Verbatim from the HealthBench paper: 1,021 physicians submitted an interest form → filtered to 683 (67%) on the quality of that form → those completed a paid introductory campaign → 268 (26%) passed and joined the HealthBench campaign → 31 more were rotated out for diversity and quality, leaving 262 (HealthBench paper). All 262 are named individually in the acknowledgements. "All members of the physician cohort were compensated for their contributions." No rate disclosed. No data-labelling company identified anywhere.
HealthBench Professional repeated the shape: 190 physicians, 50 countries, 26 specialties, selected "through a multi-step process that emphasized the quality of their application materials and paid introductory tasks", with three-stage adjudication. Same compensation sentence, same silence on rate, same absence of a vendor (OpenAI PDF).
That is a first-party funnel — inbound form, paid screening tournament, in-house quality bar, in-house adjudication ladder — run twice, by the buyer, at the exact scale a vendor would have sold.
The counter-reading is not weak. OpenAI's own Human Data postings describe a Program Manager who "will be a key interface between our external vendors and AI trainers, ensuring human data campaigns are successfully completed", and a Research Program Manager, Human Data Campaigns who must "advise and empower program managers and vendors to drive day-to-day execution" (Ashby). Campaign is the same word the HealthBench methods section uses. OpenAI demonstrably uses vendors at programme level; whether one sat behind HealthBench's tooling, screening or payments is unresolved.
Both readings survive the same facts. Bearish: a lab that has run this twice and named its 262 physicians has internalised the capability. Bullish: "vendor" and "first-party" are not opposites — the lab owns the design, the bar and the relationship, and buys operational scaffolding, which is a smaller and entirely real business. The labs as health buyers sets both out; it is the top item at What the health dossier could not establish.
The size problem
Build the cost from OpenAI's own disclosed volumes, the only auditable quantities in the sector. 700,000 model responses reviewed cumulatively by 260+ physicians (OpenAI, June 2026) against 600,000 in January (OpenAI, January 2026) — about 20,000 a month. At six minutes each, 70,000 physician-hours. Add HealthBench's 5,000 conversations and 48,562 rubric criteria (3,750 hours), HealthBench Professional's 525 adjudicated tasks (1,600), 3,500 physician-written baselines (1,750) and the 683-physician introductory campaign (2,050). About 79,000 physician-hours.
At $150/hour — Mercor's published band for internal medicine, EM and cardiology is $130–180 — that is ≈$12M. At $100/hr and four-minute reviews, $5.5M; at $200/hr and twelve-minute reviews, $30M [INFERENCE — the hour assumptions are mine, the volumes are OpenAI's]. So $5–30M cumulative, central ≈$12M, for the most physician-intensive lab programme in existence. At roughly $8M a year that is the same order as three senior health research packages at OpenAI's published $295K–$555K band once equity is counted [INFERENCE — equity is not published] (Ashby). The labs are not underspending because they cannot find suppliers. Rubric-graded evaluation is cheap relative to training.
The buyer count is smaller than the four-lab shorthand implies. Meta's careers site returns nothing clinical — every "health" match is Environmental Health & Safety or data-centre facilities (metacareers). xAI has 236 open roles, a Human Data function, and no medical vertical (x.ai). Two of six labs have no health programme at all. OpenAI and Microsoft AI are deep; Anthropic is connector-shaped with no physician cohort named anywhere; Google DeepMind is the most clinician-labour-intensive of all and describes none of it. The lab TAM is four, realistically two and a half.
Now the number that changes the frame. Pharmacovigilance outsourcing was $8.51B in 2026 (TBRC), or $9.15B on Mordor's count at 15.6% CAGR against Accenture, IQVIA, ICON and Oracle (Mordor); medical affairs outsourcing $3.05B in 2026 (MRF) or $2.59B in 2025 (Cervicorn). Roughly $11–12B of clinician judgement already priced, adjudicated and audited under regulatory scrutiny.
That is the real comparable and the real competitive set. It is not an AI market and must not be counted as TAM. What it proves is what the $12M figure hides: case-level review with blinded adjudication and named qualified reviewers is already an eleven-figure industry. You are not inventing a way to organise clinicians; you are pointing an existing one at a new buyer. See Pharma, payers and providers.
You are not first
BioStack Platforms was founded in October 2025, is in the Y Combinator Spring 2026 batch, employs seven people, and turns longitudinal clinical records into RL environments where labs train and evaluate medical decision-making — tasks, reward functions and benchmarks scored against outcomes, guidelines and clinician review. It reports 17 customers and prospects including "top AI labs" and references delivering data under a six-figure contract (YC, launch).
Seven people, ten months, six-figure lab contracts, and almost nobody has noticed. That is the strongest existence proof in the dossier and the clearest warning at once. BioStack and Sepal covers it alongside Mercor's February 2026 acquisition of Sepal AI and its 20,000+ domain experts. Mercor buys environment-construction teams, not expert networks, because it already has the experts.
The rest of the field is layered: records is a separate industry with separate buyers (The real-world data brokers); volume labelling belongs to Centaur.ai at an implied $10–25/hr (Centaur.ai, in full); and clinician access belongs to Sermo, M3 and Doximity, who own the hardest asset to build — Sermo paid $20M to members in one year — and do not yet know they are in this market (The physician panels).
One sharp negative across all of them: only Mercor has verified, named frontier-lab customers. Every other company in this vertical claims lab relationships without naming one.
The regulatory point that makes this different from security
Security has an anchor buyer that pauses training runs and publishes its evaluation vendors. Health has something structurally better and commercially slower: a regulator that wrote the product specification.
FDA's January 2025 draft guidance asks device sponsors to provide, as marketing-submission content: "a description of the expertise of those performing the data annotation"; the instructions given, "including whether annotators are blinded to each other"; the methods for "adjudicating disagreements"; and "an assessment of the intra- and/or inter-clinician variability for each task… as well as an assessment on whether the observed variability is within commonly accepted standards". It recommends "the use of independent assessments by each annotator, without knowledge of the other annotators' decisions" (FDA PDF, docket FDA-2024-D-4488).
Read as a specification, that is eight deliverables. Seven can be assembled after the fact by a diligent regulatory writer. Item four cannot. If annotators were not independent and mutually blinded at collection time, the inter-rater statistic does not exist and cannot be manufactured.
That is the moat, and it is unusually literal: the audit trail must be produced at annotation time, by a party that can attest to it (The regulator wrote your product spec). Europe reinforces it — AI Act Article 10(5), preserved and extended by Regulation (EU) 2026/1744, permits real special-category data only where the job "cannot be effectively fulfilled by processing other data, including synthetic or anonymised data" (Europe prefers the design you were going to build).
The scope limit belongs in the same breath: this binds AI-enabled medical devices, not general-purpose LLMs deployed as clinical assistants — which is exactly where the lab money sits. The regulator's demand and the labs' budget point at two different buyers.
The six scores
| Axis | Screen: clinical | Deep look | Security |
|---|---|---|---|
| proof | 4 | 4 | 4 |
| budget | 5 | 4 | 5 |
| defense | 3 | 4 | 3 |
| cost | 2 | 4 | 3 |
| reach | 4 | 5 | 4 |
| room | 3 | 3 | 3 |
| total | 21 | 24 | 22 |
proof: 4. The scoreboards are public and the buyers are losing on them: Vals MedCode at 63.57% for the best model on earth across 86 models (Vals), VariantBench with no model-harness pair above 50% (latch.bio), MedHELM's Administration & Workflow at 0.53–0.63, the lowest band in a 121-task taxonomy (arXiv), and nursing with nothing at all. BELO even itemises the recipe — 900 questions, ~23 ophthalmology professionals, three review tiers, senior adjudication (arXiv). Not a 5, because proof means a reference customer who says so out loud and no company in this vertical has ever got a lab to say its name.
budget: 4, down from 5. Money moves: 262 then 190 compensated physicians at OpenAI; eighteen named healthcare roles with bands on Mercor's page; 7,700 US-licensed clinicians hired by Hippocratic AI for 775,000 test calls (Hippocratic); Ambience benchmarking its ICD-10 model against 18 board-certified physicians with four-or-more-clinician consensus per encounter; BioStack's six-figure contract. Against it: ≈$12M central at the deepest lab, two of six labs absent, pharma buying platforms not labels in every disclosed deal 2020–2026, no payer found buying anywhere. Nothing here matches security's Astra event.
defense: 4, up from 3. The credential is registry-verifiable and cannot be faked, which is not true of "senior security practitioner". The content churns on a published schedule — ICD-10-CM every October, CPT every January, Coding Clinic and NCCI quarterly — so an evaluation set decays by construction and a subscription is the natural contract. And item four permanently favours whoever collected blinded. Not a 5 because in the highest-revenue niche the barrier is a mailing list, against 866,460 physicians in direct patient care (AAMC).
cost: 4, up hard from 2 — the largest correction on the page. No rigs, no scanners, no bare-metal estate, no data-use agreement; the PHI-free architecture removes BAAs, expert determinations, Security Rule scope and breach exposure outright (Build it without ever touching a patient record). The first coding pilot is twenty coders, a spec, a dual-coding harness and an adjudicator. Held off 5 by the product itself: three blinded readers per item costs three times a single label.
reach: 5, the best-evidenced 5 in the atlas, because in healthcare the state maintains your register. NPPES NPI is a free, public, API-accessible national provider database (CMS) — and Anthropic shipped an NPI connector in January 2026, so the labs already treat it as the canonical clinician identity primitive. Add Nursys plus fifty state boards, NABP Verify, AAVSB's VAULT, AHIMA's date-stamped counts (CCS 36,925, CCA 7,753, RHIT 26,128, CDIP 2,913 at 31 December 2025, AHIMA), and near-census societies — AAO covering >90% of practising US ophthalmologists [WEAK — Wikipedia, not primary], ASCO >50,000 across >150 countries (ASCO), AVMA's 133,475 US veterinarians (AVMA), AAPC's 300,000+. The caveat is conversion: emailing 12,000 physicians from IQVIA's OneKey database converted 1.2% to 149 completers (arXiv).
room: 3. BioStack is executing the thesis; Mercor runs a full healthcare vertical; Centaur.ai holds volume labelling; mpathic holds behavioural health with $15M and a published benchmark; Vals and Protege hold evaluation benchmarking; Hippocratic internalised nursing. Off a 2 because nobody occupies the FDA-conformant blinded-panel position, nobody sells characterised disagreement as a product, and four benchmark landscapes — nursing, deprescribing, pharmacovigilance, veterinary — are empty.
24 against the screen's 21. The deep look lowered the budget and raised almost everything else, because the screen priced the wrong labour.
The thirteen, ranked
Composite sums the six axes on each niche page. It sorts; it does not decide.
| # | Niche | Composite | Stance | The one sentence |
|---|---|---|---|---|
| 1 | Medical coding | 26 | build | 3x spread, no licensure, no PHI, best model at 63.57% |
| 2 | Dermatology and ophthalmology | 24 | build | The only reimbursed autonomous AI service, CPT 92229; AAO is a census |
| 3 | Nursing | 24 | build | 3.47M verifiable RNs, no benchmark at all, hostile union |
| 4 | Genomics and variant curation | 23 | watch | Cleanest artefact anywhere, ~440,000 unresolved VUSs, no buyer |
| 5 | Veterinary medicine | 23 | watch | No HIPAA; the real play is toxicologic pathology, i.e. pharma money |
| 6 | Clinical reasoning | 22 | build | The only proven repeat buyer, and four incumbents already selling |
| 7 | Pharmacy | 22 | watch | Everything works except the 1.1–1.8x multiple and the absent buyer |
| 8 | Pathology | 21 | watch | Ground truth genuinely is a distribution; ~90% of US slides are still glass |
| 9 | Behavioural and mental health | 21 | watch | Highest willingness-to-pay, occupied by mpathic, heaviest safety load |
| 10 | Clinical trials and regulatory writing | 18 | avoid | Pharma buys platforms; freelance writers already earn $131.65/hr |
| 11 | Oncology | 18 | avoid | NCCN owns the guidelines and licenses nothing |
| 12 | Radiology | 17 | avoid | The arbitrage runs backwards |
| 13 | Ambient scribing | 16 | avoid | MedScribe is at 90.99% and the raters are physician-priced |
Two places the composite misleads. Clinical reasoning ranks sixth and should be entered first — the only niche with disclosed, repeated, admitted lab purchasing. Enter it for revenue, knowing the moat is a mailing list, and let it fund the defensible work. And Genomics and variant curation and Pharmacy score high on axes that quietly assume a buyer appears; both score 1 and 2 on budget because nobody has been found paying. High composite with low budget is a solution looking for a customer.
When the answer flips
To no. If a named vendor did run HealthBench's tooling and payments and labs simply do not credit suppliers, the first-party argument inverts — but so does the opportunity, because that vendor is the incumbent with two years of relationship on you. If BioStack converts its 17 prospects and raises on named lab logos, room drops to 2 and the move is to sell into it. If Sermo, M3 or Doximity announces an AI-data line, the supply argument dies the day it ships. And if buyers will not pay for three blinded readers where one label is cheaper, this is a single-annotator labelling business — which is the trap.
To a much louder yes. If any lab or device sponsor confirms buying clinical expert data from a vendor — nothing in the record says yes — the category converts from inference to arithmetic. If FDA finalises the January 2025 draft with the blinding and variability language intact, item four becomes submission content, and the installed base is large: 1,016 authorisations through December 2024 across 736 unique devices (npj Digital Medicine). If a specialty society will co-sign a methodology — and no medical society has taken any position on clinicians selling judgement for AI training, a void rather than a rejection (What the doctor on the other end is risking) — recruitment and procurement are solved at once. And if MedCode stays under 70% through 2027 while MedScribe sits at 91%, the coding thesis has an eighteen-month clock, not a six-month one.
The specific opportunity is at Characterised disagreement. What a clinician hour costs is at What a clinician hour costs. The test sequence is at Ninety days in health, and what could not be established is at What the health dossier could not establish.