Every conversation about this business eventually collapses into one question: what is the unit, and what does it cost to make? The public record contains exactly one clean answer.
Mercor built its APEX benchmark — 200 tasks across law, medicine, finance and management consulting — for over $500,000, with advisors from Harvard Business School and Harvard Law School (Time). That implies roughly $2,500 of expert labour per benchmark task.
$2,500 per benchmark task, from a $500K spend on 200 tasks across four professions. At a $110–250/hr physician rate, that is 10–23 hours of expert time per task — authorship, verification, adjudication and revision included. If you are pricing a clinical evaluation suite, this is the closest thing to a market comparable that exists.
Set that against what the labs actually buy at. OpenAI's entire physician programme — 700,000 reviewed responses, 48,562 rubric criteria, two published benchmarks — plausibly cost $5–30M in total, central estimate ~$12M (The labs as health buyers). A 200-task benchmark at $2,500 a task is $500K. The gap between those numbers is the whole commercial question: is the sellable unit a benchmark, or is it a subscription to continuous judgement?
The three-party value chain that has already assembled itself
The most instructive thing in the competitive record is not a company. It is a structure that formed without anyone designing it.
In February 2026, Protege supplied the clinical data behind two new Vals AI healthcare benchmarks. The graders were certified professional coders. The academic imprimatur came from Professor Pranav Rajpurkar at Harvard Medical School (Protege, Vals).
The construction details are the useful part:
- MedScribe — SOAP-note generation from doctor–patient transcripts. 80 transcripts, 100 expert-developed assessment criteria, every sample independently annotated. Transcripts were synthetic, generated from real de-identified SOAP notes via an adapted NoteChat framework; realism was validated by an A/B test in which medical experts identified the true transcript only 39% of the time (chance = 50%). Rubrics were authored by medical scribes and physician assistants who annotated independently, with conflicts reconciled collaboratively (Vals).
- MedCode — ICD-10-CM assignment from de-identified discharge summaries with supporting progress notes. 2,755 samples, each independently annotated by two certified coders using professional tooling, with patient-level holdouts to prevent memorisation (Vals).
So the emerging structure of clinical evaluation is: a data licensor + a benchmark house + credentialed graders + an academic imprimatur. Four legs, three companies, one professor, assembled ad hoc.
Protege has raised ~$65M ($25M Series A August 2025, $30M extension led by a16z January 2026) and holds 3B+ clinical notes and 100M medical images, paying data holders through revenue share on each use (Protege). Vals AI raised a $40M Series A at a $400M valuation on 13 August 2026, led by a16z, on top of a $5M seed (Sacra, Teknowire).
Owning all three legs is the differentiated position. Not because integration is elegant, but because each leg alone is weak. A data licensor is a records business with the margin structure of The real-world data brokers. A benchmark house without proprietary data grades whatever it is handed, and Sacra notes Vals already flags expert-sourced benchmark creation as a margin drag. Credentialed graders alone are a staffing agency competing with the panels on wage. The combination is the only configuration in which the rubric — the actual instrument — is proprietary, and the instrument is what a lab cannot easily rebuild. Vals's own economics make the argument: Sacra's read is that public benchmark investment is not a marketing cost, it is the mechanism that makes private enterprise evals credible enough to sell. That is Publishing the benchmark with a P&L attached.
The sellable units, from cheapest to most defensible
| Unit | What it is | Who buys | Price evidence |
|---|---|---|---|
| Labels | A classification on an existing record | Device sponsors, imaging AI, health systems | ~$10–25/hr crowd equivalent (Centaur.ai, in full) |
| Rubric criteria | Written pass/fail conditions a response must satisfy | Labs, benchmark houses | 48,562 criteria in HealthBench; no per-criterion price published |
| Adjudicated disagreements | Two independent reads plus a senior tiebreak | FDA-facing sponsors, labs | None public; the FDA asks for it by name |
| Reasoning traces | The clinician's working, not just the answer | Labs, post-training teams | None public |
| De-novo cases | Clinician-authored synthetic case with findings and plan | Labs, benchmark builders | None public; the PHI-free route |
| Held-out evaluation sets | A benchmark whose answers are not public | Labs, third-party evaluators | ~$2,500/task (APEX) |
| RL environments | A scored, interactive clinical task | Labs | "Six-figure contract" (BioStack and Sepal) |
| Reward models | A trained grader that scores without a human | Labs | None public |
The ordering is not arbitrary. It runs from what an offshore documentation specialist can produce — Commure does exactly this in Bengaluru (Health-AI companies as buyers) — to what requires a specialty-matched physician, an adjudication protocol and an engineer who can build a scored environment. Price defensibility tracks that ordering exactly, and so does the difficulty of building it.
Three units deserve singling out.
Adjudicated disagreements are the unit regulation asks for by name. The FDA's January 2025 draft guidance requires marketing submissions to include "a description of the expertise of those performing the data annotation," the "number of participating clinicians and their qualifications," "the methods for evaluating quality/consistency of data annotations and adjudicating disagreements," and "an assessment of the intra- and/or inter-clinician variability for each task" — and recommends "the use of independent assessments by each annotator, without knowledge of the other annotators' decisions" (FDA). That converts blinding and adjudication from good practice into submission content. It is the strongest structural argument in the dossier and it is scoped to AI-enabled devices, not to general-purpose LLMs — see The regulator wrote your product spec.
De-novo cases are the PHI-free route, and the one OpenAI's physician cohorts actually took. A clinician writes the case; no hospital, no IRB, no de-identification pipeline, no institutional data ownership. It is also the route with the lowest technical moat, precisely because it needs none of that machinery — Build it without ever touching a patient record.
RL environments and reward models are where the acquisition record points. Mercor bought Sepal and Deeptune for environment-construction capability, not for expert networks, and BioStack's most expensive open role is an RL research engineer at $200–350K rather than its Clinical Data Lead at $120–200K. Every buyer in this market is signalling that the scarce input is the person who builds the scored task, not the person who answers it.
What the annotated-data price actually is
One transaction puts a number on annotated specialist data. Tempus acquired Paige for $81.25M in August 2025, explicitly for its "proprietary dataset of almost 7 million digitized pathology slides that are clinically annotated" (Tempus IR, Fierce Biotech).
Naively that is under $12 per annotated slide, and the naive reading is wrong in both directions: the price bought a company, its staff, its FDA clearances and its software as well as the corpus, and the corpus was accumulated over a decade of clinical operation rather than commissioned. Still, it is the nearest thing to a market clearing price for annotated specialist medical data at scale, and it is a low number — an acquirer paid $81M for seven million expert-touched artefacts. Anyone modelling a per-label business against a $2,500-per-task benchmark comparable should hold both numbers at once. They describe opposite ends of the same market.
The other comparable in the frame is OpenAI's Torch acquisition, reported at $60M (Becker's) or ~$100M in equity (TechCrunch) for a four-person health-records team. Terms were never officially disclosed; both figures are reported, not confirmed. That is infrastructure, not judgement — but it establishes that a lab will pay eight figures for health data plumbing while spending perhaps $12M in total on physician judgement.
No vendor publishes a rate card
State it plainly, because it shapes every negotiation. No health data-labelling or expert-data company publishes a rate card. Not Centaur, not Protege, not Vals, not BioStack, not iMerit, not Shaip. Vals sells through demo-led negotiated contracts blending platform subscription with usage-based evaluation volume, and publishes no pricing at all (Sacra).
The three usable public anchors, in full:
- ~$2,500 per benchmark task (Mercor APEX, 200 tasks, >$500K)
- "Six-figure contract" (BioStack, one reference, no term)
- Six- to seven-figures annually for national claims feeds (Alpha Sophia) — a records comparable, not a judgement one
There is no per-rubric, per-adjudicated-disagreement or per-reasoning-trace price anywhere in public — which is the unit a clinical judgement business would actually sell. HealthBench's 48,562 criteria from 262 physicians implies about 185 criteria per physician, with no time and no cost attached. Every financial model in this market is extrapolating from hourly rates and headcounts, not from observed deal sizes. Nobody has published a contract.
The unit is not the label and it is not the hour. It is the instrument — a rubric, an adjudication protocol, a scored environment — and the one public price for building one is $2,500 a task. The value chain that produces those instruments today runs across three companies and a professor. Owning all three legs is the differentiated position, and the reason is margin, not tidiness: it is the only configuration in which the thing you sell cannot be reassembled by the buyer from parts. Read with The measurement landscape and Characterised disagreement.