Every other page in this dossier is trying to assemble the demand proof that clinical reasoning already has. OpenAI built HealthBench with 262 physicians from 60 countries across 26 specialties, producing 5,000 conversations and 48,562 unique rubric criteria (OpenAI). It then built HealthBench Professional with 190 physicians from 50 countries, distilling 15,079 candidate examples down to 525, and stated in the paper that "all members of the physician cohort were compensated for their contributions" (HealthBench Professional). ChatGPT for Clinicians shipped on 22 April 2026 built with "hundreds of physician advisors" who tested 6,924 conversations in daily work (OpenAI).
That is three separate paid physician cohorts from one buyer, disclosed in writing, inside eighteen months. Nothing else in healthcare comes close.
The catch is in the same sentence. Every artefact in those programmes is text — written cases, rubrics, ratings of model outputs. Text needs no scanner, no DICOM pipeline, no institutional agreement and no fellowship. Which is precisely why Mercor, Handshake, Surge and Outlier are all already selling it.
What the artefact is
Six things, all writable at a keyboard:
- Differential-diagnosis construction — ranked hypotheses with the explicit discriminator for each.
- Workup planning — which test next, why, at what cost.
- Reasoning traces on hard cases — the token-level "thinking" that RL post-training consumes.
- Multi-turn agentic episodes where the model must ask for information rather than receive a tidy case summary. Microsoft's SDBench is the canonical shape here: 304 NEJM clinicopathological conference cases restructured so the model interrogates a gatekeeper agent (arXiv 2506.22405).
- Diagnostic-error adjudication — retrospective analysis of a miss with the cognitive-error taxonomy attached.
- De-novo hard cases, written by subspecialists to be deliberately adversarial.
Mercor's own listing describes the job in almost these words: "write gold-standard clinical responses, realistic patient cases and evaluation rubrics," "design challenging scenarios that reveal how models perform" (Mercor). The artefact is not a secret. It is a job advert.
The market it addresses is not small. Diagnostic error causes an estimated 795,000 Americans a year to suffer death or permanent disability — 371,000 deaths and 424,000 permanent disabilities — with the "Big Three" of vascular events, infections and cancers accounting for 75% of serious harms and an average missed-diagnosis rate of 11.1%, from 1.5% for heart attack to 62% for spinal abscess (Johns Hopkins).
Does it need patient data
No. This is the cleanest PHI position of the six, and it is the reason this is the niche the labs actually funded.
A physician writing a case from professional knowledge rather than from a chart creates no protected health information, no human subject and no institutional data-ownership claim. There is no imaging, no DICOM header, no burned-in pixel text, no IRB. NEJM clinicopathological conferences are already published. The de-novo route here is not a workaround — it is the product, and it is exactly what the HealthBench cohorts did. See Build it without ever touching a patient record for the full argument, including the threshold point that a European operator paying clinicians to write cases is neither a covered entity nor a business associate and therefore sits outside HIPAA's jurisdiction rather than merely complying with it.
A clinician rating a model's output about a hypothetical patient is likewise clean on every axis: no individually identifiable health information exists, no living individual is the subject, and the rating is an opinion about a text rather than about a person.
The same fact is the strategic problem. Zero PHI means zero technical moat. There is nothing to acquire, license, de-identify or host. The only asset is the roster.
Is anyone buying
Yes, and budget scores 5 — the only uncontested 5 in the health dossier.
Beyond OpenAI's three cohorts, the buyer set widens. Microsoft's MAI-DxO reached 80% accuracy, 85.5% at best configuration, against 20% for 21 generalist physicians on SDBench (arXiv). Google DeepMind's AMIE line has published twice in Nature and is the most clinician-labour-intensive research programme at any lab — OSCE-style evaluation needs simulated-patient actors, physician participants and physician graders, three separate paid populations. Stanford's MedHELM taxonomy was developed and validated by 29 clinicians across 121 tasks (arXiv 2505.23802).
And the retention is now visible. OpenAI describes "a global network of more than 260 physicians" who have "reviewed more than 700,000 example model responses," against 600,000 five months earlier — a run-rate of roughly 20,000 reviews a month (OpenAI). That is not a one-off benchmark build. That is a standing panel with a burn rate. See The labs as health buyers.
Proof scores 3, not higher. Demonstrating you are the best source of clinical reasoning data is hard precisely because the artefact is legible and four funded companies are already producing it. You cannot be first; you can only be better, and "better" here means harder cases and tighter adjudication, which takes a publication cycle to show.
What the expert costs
Cost scores 5 — the cheapest pilot in the dossier. A first cohort is twenty physicians, a rubric spec and a shared document. No scanner, no licence fee, no data-use agreement, no hospital BD.
The labour is not cheap in absolute terms but it is the cheapest relevant labour: blended cost sits between internal medicine (~$307K, roughly $154/hr at 2,000 hours [INFERENCE]) and the diagnostic subspecialties at $394–571K. Observed AI-data rates run $110–250/hr on Mercor, against a national AI-trainer average of $31/hr (Mercor).
The margin does not come from paying doctors less than their day job — see What a clinician hour costs. It comes from selling marginal hours they would otherwise not work, and from the adjudication wrapper around them.
Getting to them
Reach scores 5. The addressable pool is essentially every practising physician: 1,032,365 active US physicians, 866,460 in direct patient care (AAMC). That is simultaneously the strength and the weakness — huge supply means no scarcity, and no scarcity means no pricing power.
Named channels: r/medicine at 546,345 subscribers (RedPulse) [WEAK — third-party tracker]; academic hospitalist and diagnostic-excellence communities; the Society to Improve Diagnosis in Medicine; and the NPPES NPI Registry, a free public API covering every US provider (CMS). NPI gives identity and claimed specialty only — taxonomy codes are self-attested — so it is step one of a four-step stack ending in board-certification lookup.
Where the benchmarks sit
| Benchmark | Status | Frontier score |
|---|---|---|
| HealthBench Hard (1,000 examples) | Far from saturated | gpt-5-thinking 46.2%, o3 31.6%, GPT-4o 0.0% |
| HealthBench Consensus | Near-saturated | gpt-5-thinking-mini 96.5% |
| HealthBench Professional (525) | Open, new | ChatGPT for Clinicians 59.0, base 48.1, human physicians 43.7 |
| SDBench (304 NEJM CPC) | Open | MAI-DxO 80% vs 21 physicians at 20% |
| MedXpertQA | Contested | top 0.784 (Muse Spark), August 2026 |
| MedHELM — Clinical Decision Support | Weakest category in the suite | 0.56–0.72 across nine frontier models |
Scores from the GPT-5 system card, the HealthBench Professional PDF, SDBench, llm-stats and MedHELM. More at The measurement landscape.
Read two things off that table. HealthBench Hard at 46.2% and MedHELM Clinical Decision Support at 0.56–0.72 mean the open capability gap sits exactly where this product sits. And models already beat physicians on HealthBench Professional, 59.0 to 43.7 — which is the uncomfortable half.
What would kill it
Commoditisation, and it is already happening. Defense scores 2 and room scores 2.
Mercor posts $110–250/hr for this work today. Handshake AI advertises physicians at up to $200/hr "supporting AI research for companies like OpenAI and Anthropic" [WEAK — single aggregator]. Outlier runs a radiology vertical whose described tasks are text QA and which explicitly says "clinically experienced NPs, PAs, or nursing professionals may also be considered" (Outlier) — a cheap product sold as specialist expertise. With 866,460 physicians in the frame and no credential barrier beyond a licence check, supply cannot be cornered. See Expert data for frontier labs on why the generalist marketplaces win commodity artefacts.
The second kill is subtler. If the model already scores 59.0 against the physician's 43.7, the commercial logic of buying more expert labels shifts from producing ground truth to producing harder cases and disagreement signal. That is a different product with a different price, and nobody has published evidence that buyers have made the shift. Characterised disagreement takes this up.
The third is displacement by the buyer. OpenAI's Human Data function runs the vendor interface at programme level with three open roles paying up to $385K; a lab that assembles its own 260-physician standing panel does not need yours.
Across the whole diagnostic side of healthcare, there is not one disclosed contract between a frontier lab and a specialist-physician data vendor. The evidence here — the strongest in the dossier — is still inferential: benchmark acknowledgements that physicians "were compensated," platform rate cards, and connector announcements. No dollar values, no term lengths, no named vendor-lab pairs. Anyone modelling this business is extrapolating from hourly rates and headcounts, not from observed deal sizes.
Where the record is thin
There is no per-artefact price anywhere. We have per-hour expert rates and per-image annotation prices; we have nothing on what a lab pays per rubric, per adjudicated disagreement or per reasoning trace — the unit that would actually be invoiced. HealthBench's 48,562 criteria from 262 physicians implies roughly 185 criteria each, with no time or cost attached [INFERENCE, weak].
Nor is it established whether OpenAI's 260-physician network is sourced through a vendor at all, or recruited directly. The Program Manager posting confirms external vendors are used for human data campaigns generally; no vendor is named on any health artefact.
Compare against the credentialled-professions view in Clinical medicine, and the ranked read in The health read.