Miju Labs

The security dossier

The labs as health buyers

OpenAI published its physician recruitment funnel in full and no vendor appears anywhere in it — 1,021 applicants down to 262 paid physicians, run first-party. The lab TAM is four, realistically two and a half, and the whole programme has plausibly cost $12M.

high confidence10 minupdated 2026-08-30frontier labs · health · buyers · healthbench · hiring

The question this dossier turned on was whether a frontier lab buys clinician judgement from a vendor or assembles it itself. OpenAI has answered it in writing, twice, and the answer is: itself.

The HealthBench paper discloses the entire pipeline. 1,021 physicians submitted an interest form. OpenAI filtered to 683 (67%) "based on needs of the campaign and the quality of their responses to the interest form." Those physicians completed a paid introductory campaign of tasks resembling the real work. 268 (26%) passed and joined the HealthBench campaign; 31 were later removed for diversity and quality; 262 are named individually in the acknowledgements (HealthBench paper). Eligibility was medical school completed and active practice within the last five years. On money, one sentence: "All members of the physician cohort were compensated for their contributions." No rate. No data company named anywhere in the paper.

HealthBench Professional repeated the shape a year later with 190 physicians across 50 countries and 26 specialties, selected "through a multi-step process that emphasized the quality of their application materials and paid introductory tasks," with automated quality checks, manual rubric review and three-stage adjudication — and the identical compensation sentence, again with no rate and no vendor (OpenAI PDF).

The finding

A frontier lab ran its own inbound interest form, its own paid screening tournament and its own adjudication tiers to assemble 262 physicians across 60 countries — and named no intermediary. That is either the strongest evidence of unmet demand in this dossier, or the strongest evidence that labs prefer to do this in-house. It cannot be both, and which one it is decides the business.

The counterweight: vendors exist, one layer up

OpenAI's own Human Data job postings cut against the pure first-party reading. The Program Manager, Human Data "will be a key interface between our external vendors and AI trainers, ensuring human data campaigns are successfully completed," and the Research Program Manager, Human Data Campaigns is tasked to "advise and empower program managers and vendors to drive day-to-day execution" (OpenAI Ashby board).

The word that matters is campaign — the same word HealthBench's methods section uses for "the introductory campaign" and "the HealthBench campaign." So vendors are in the loop at the programme level, for human data generally, even where none is named on a specific health artefact.

The honest reconstruction: the lab owns the design, the quality bar, the adjudication and the physician relationship; vendors supply operational scaffolding — tooling, payments, throughput. Whether one sat behind HealthBench specifically is unverified either way, and the paper's own phrasing ("physicians expressed interest in working with us") points first-party. That division of labour is not a rounding error for a would-be supplier: scaffolding is a low-margin, replaceable position, and it is the only slot the published record confirms is open. See What actually gets sold for what sits on the other side of that line.

Only two roles in the world require a medical degree

Across every lab board readable on 30 August 2026, exactly two open roles require a medical licence: Anthropic's Partner Manager, Global Health at $215K–$300K OTE, demanding "Medical training and clinical practice (MD, GP, MBBS, DO, or equivalent)" (Greenhouse), and Microsoft AI's Clinical Specialist in London at £93,500–£161,800, demanding a medical degree with two years' postgraduate clinical experience (microsoft.ai). Google's Clinical Specialist, Health Optimization in New York requires a doctoral clinical degree but its band renders client-side and could not be retrieved.

OpenAI has zero licence-requiring roles. Its top health band is a machine-learning job: Research Engineer / Research Scientist, Health at $295K–$555K, sitting in Post-training inside the Personal AGI organisation (Ashby). The lab with the deepest physician programme in the industry employs no physicians on its own payroll to run it. It rents them, project by project, off an inbound funnel — which is exactly the shape of demand a supplier can serve, and exactly the shape of demand a supplier can be cut out of.

Microsoft AI is hiring your buyer

The single most commercially legible posting in the set is Microsoft AI's Sr. Director, AI Data Acquisition and Operations at $202.4K–$303.6K in the Bay Area. It "partners with researchers to identify critical data needs, select vendors, and personally negotiate commercial agreements," oversees "human data and annotation vendor programs," manages "spend tracking, vendor payments, and milestone delivery" — and is charged with "structuring novel deals for data that have never been commercially available before" (microsoft.ai).

That last clause is close to a job description for the counterparty. Microsoft AI also carries 13+ open health roles, more than any other lab, including a Member of Technical Staff — AI Evaluations, Health built around its Mayo Clinic partnership, which describes working with "Mayo clinicians as subject-matter experts" and "recruiting physician raters" (microsoft.ai). Read that alongside the buyer role and the message is mixed: Microsoft is staffing to buy, and simultaneously staffing to get physician raters free through an institutional partnership that also confers a credibility halo no vendor can match.

Two of six labs have no health programme at all

Meta: every "health" match on its careers site is environmental health and safety or data-centre facilities (metacareers.com). xAI: 236 open roles, zero medical. It runs a Human Data function and an "AI Tutor" contractor family spanning 30 languages plus software engineering and finance — and no medical specialism (x.ai). Both are negative findings from direct inspection of the boards, not absent search results.

So the addressable lab set is four, not six — and Google's health work sits mostly in Research and Platforms & Devices rather than a buying function, and Anthropic's evaluation posture is benchmark-led rather than clinician-rated (MedCalc, MedAgentBench, SpatialBench; no physician cohort named in the Claude for Healthcare launch, Anthropic). Call it two and a half. That is the concentration problem from One customer is a binary event in its most acute form: a supplier here has fewer plausible customers than a defence contractor.

Every lab health posting, with bands

Live on 30 August 2026. Roles sharing a band are grouped; "—" means no band published in the reachable source.

TitleLabBandSource
Research Engineer / Scientist, HealthOpenAI$295K–$555Kashby
SWE, Research — Human DataOpenAI$230K–$385Kashby
Full Stack SWE, Health AIOpenAI$293K–$325Kashby
Research PM, Human Data CampaignsOpenAI$239K–$328Kashby
Model Policy, Chemical & Biological RiskOpenAI$207K–$295Kashby
Program Manager, Human DataOpenAI$207K–$230Kashby
FDE, Healthcare ×3 (SF, NYC, Seattle)OpenAI$198K–$280Kashby
Account Director, Healthcare / Life Sciences ×2OpenAI$189K–$240K + commissionashby
Account Director, Large Enterprise, Life Sciences (London)OpenAIashby
Research Engineer, Life SciencesAnthropic$350K–$500Kgh
Life Sciences CounselAnthropic$335K–$385Kgh
Mgr, Applied AI Eng / Architecture, Healthcare & Life Sciences ×2Anthropic$315K–$405Kgh
RS/RE, Biological SafetyAnthropic$300K–$405Kgh
Research Scientist, Life Sciences ×3 (incl. Chemistry, Computational)Anthropic$300K–$320Kgh
Life Sciences Operator, LeadAnthropic$300K–$320Kgh
Applied AI Engineer, Beneficial DeploymentsAnthropic$280K–$320Kgh
Safeguards Enforcement Analyst, Bio HarmsAnthropic$245K–$285Kgh
Partner Manager, Global Health (MD required)Anthropic$215K–$300Kgh
Research Operations Lead, BiologyAnthropic$150K–$200Kgh
Enterprise AE, Life Sciences (London)Anthropic£280K–£330Kgh
Sr. Director, AI Data Acquisition and OperationsMicrosoft AI$202.4K–$303.6K (Bay/NYC)msft
MTS — AI Evaluations, Health (Mayo partnership)Microsoft AIIC4 $160.2K–$261.0K; IC5 $188.0K–$304.2K (Bay/NYC)msft
Clinical Specialist (medical degree required)Microsoft AI£93.5K–£161.8Kmsft
10 further MAI Health roles (product, applied AI, FDE, privacy, design, TPM, commercialisation)Microsoft AImsft
Senior Research Scientist, Foundational AI in HealthGoogle$174K–$252K + bonus + equitycareers
Clinical Specialist, Health Optimization (doctoral clinical degree)Google— (renders client-side)careers
Research Scientist, Frontier Health; regulatory, sensing, partnerships, Cloud sales ×5Google / DeepMindcareers
None foundMetametacareers
None found (236 open roles)xAIx.ai

What the whole programme has cost

No lab discloses clinical-data spend, so build it from the volumes OpenAI does publish. 700,000 model responses reviewed by the 260+ physician network as of June 2026, up from 600,000 in January — about 20,000 reviews a month (OpenAI, Jun 2026, Jan 2026). Plus 48,562 rubric criteria over 5,000 conversations, 525 HealthBench Professional tasks with three-stage adjudication, 3,500 physician-written baseline responses, 6,924 clinician-tested conversations before the ChatGPT for Clinicians launch (OpenAI), and the 683-physician paid screening round.

At six minutes a review, 45 minutes a rubric-bearing conversation, three hours a Professional task, thirty minutes a baseline and three hours of screening, that totals roughly 79,000 physician-hours. At $150/hour — the blended rate implied by Mercor's published $130–180/hr for internal medicine, EM and cardiology (Mercor) — that is ≈$12M. At $100/hr and four-minute reviews, $5.5M; at $200/hr and twelve-minute reviews, $30M.

The disciplining figure

$5–30M, central ~$12M — everything OpenAI has plausibly spent on physician judgement since early 2025, across two published benchmarks and 700,000 reviewed responses. Three Research Engineer, Health hires at the top of band plus one Health AI full-stack engineer exceed $2M a year in salary alone. The most physician-intensive lab programme in existence costs less than three of that lab's own open health job packages per year.

That is not a lab underspending for want of suppliers. It is the price of the problem as currently framed: rubric-graded evaluation is cheap next to training compute, and a 262-physician network turning over 20,000 responses a month is evidently enough to move the benchmark. Grow the number and you need a different framing of the work — harder cases, disagreement signal, environments — which is the argument in The measurement landscape and Characterised disagreement.

What the labs will not say

No lab has published a rate paid to any physician. Google DeepMind's AMIE line — three papers across 2025–2026, all resting on OSCE-style evaluation with paid patient actors and physician raters and physician graders — gives no numbers, no recruitment method and no vendor in any public blog post. It is the most clinician-labour-intensive research programme at any lab and nobody outside it knows how it is staffed.

Where this leaves the buyer map: the labs are the prestige customer and the small one. The volume, if it exists, is next door — see Health-AI companies as buyers and Pharma, payers and providers, and read both against Expert data for frontier labs and GMV is not revenue before assuming lab logos convert into revenue.