Miju Labs

The security dossier

Characterised disagreement

Models now beat physicians on HealthBench Professional, 59.0 against 43.7. That is good news, because it moves the product from producing the right answer to producing a measured map of where the right answer is contested — which is the one artefact that appreciates as models improve, and the one FDA asks for by name.

medium confidence9 minupdated 2026-08-30product · reference standard · inter-rater · pricing · panels

The expert is no longer the ground truth. On HealthBench Professional, ChatGPT for Clinicians scores 59.0 against human physicians' 43.7 on the same 525 physician-authored tasks (OpenAI). On SDBench — 304 NEJM clinicopathological cases restructured so the model must interrogate a gatekeeper for findings — MAI-DxO reaches 80%, and 85.5% in its best configuration, against 21 practising physicians at 20% (arXiv 2506.22405).

The instinctive reading is that this ends the business. If the model outperforms the doctor, why buy the doctor?

The correct reading is the opposite, and it is the whole thesis. When the model exceeds the reference standard, measured error stops being a property of the model and becomes a property of the labels. You can no longer distinguish "the model is wrong" from "the reference is wrong". The buyer's problem changes from get me the right answer to tell me how far from truth my reference standard is, and where. Nobody sells that. It is the product.

Ground truth is a distribution, and pathology proves it

Eight board-certified dermatopathologists, median thirty-four years' experience, independently reviewed 792 melanocytic lesion slides from 736 patients across eight German university hospitals. Complete eight-of-eight agreement in only 53.5% of cases. No majority at all in 9.1%. Overall Fleiss' κ = 0.701, but only κ = 0.428 for non-invasive melanomas — the cases that actually decide whether someone gets an excision. The expert panel disagreed with the original local pathologist in 14.9% of cases (Nature Communications, 17 January 2025).

Read that as a data problem rather than a medical one. For nearly half of these slides there is no such thing as the label. There is a distribution over expert opinion, and any single number you write in the label column is a draw from it. The same is documented, less precisely, for PD-L1 CPS scoring in upper-GI adenocarcinoma (Modern Pathology) and HER2-low IHC (PubMed 36856777), where exact κ values could not be retrieved [UNVERIFIED] but the irreproducibility is the published finding.

A single opinion has no error structure. A blinded panel does.

One annotator gives you a value. Three independent, mutually blinded annotators give you a value and a variance, a disagreement map, an adjudication trail, and a per-task statement of which cases are decidable. The first is a row in a table. The second is a measurement instrument with a stated precision — and precision is the thing a regulated buyer cannot manufacture, borrow or infer.

FDA asks for exactly this, in writing

The January 2025 draft guidance does not assume the expert is truth. It asks sponsors to quantify how far from truth the expert is. Among the required contents of a marketing submission:

In their words

"A description of the uncertainty inherent in the selected reference standard." … "An assessment of the intra- and/or inter-clinician variability for each task, as applicable, as well as an assessment on whether the observed variability is within commonly accepted standards for a particular measurement task." … "FDA recommends the use of independent assessments by each annotator, without knowledge of the other annotators' decisions."

FDA draft guidance, docket FDA-2024-D-4488

The regulator has already conceded the premise this page is built on. It is not asking for expert labels; it is asking for characterised expert labels — labels with a known error structure, produced by a named, credentialed, mutually blinded panel with a documented adjudication procedure. The full reading is at The regulator wrote your product spec; Europe arrives at the same place from a different direction at Europe prefers the design you were going to build.

And the asymmetry is permanent. Seven of the eight things FDA asks for can be written up afterwards by a competent regulatory writer. The inter-rater statistic cannot. If the readers were not independent and blinded when they read, the number does not exist and no amount of later effort creates it. Whoever collected blinded owns something whoever collected consensus-first can never retrofit.

Why this appreciates rather than depreciates

Every other expert-data product in this atlas gets less valuable as models improve. This one gets more valuable, for a mechanical reason.

As a model's accuracy rises, the fraction of its residual error that is genuine model error falls, and the fraction that is label noise rises. Past some performance level the measurement is dominated by the reference standard's own uncertainty — the label-noise literature has documented the mechanism for years, and documented that the noise is class-dependent, so it biases as well as blurs (Medical Image Analysis 2020, Proc SPIE 2023). A 2026 paper in Ultrasound in Medicine & Biology is titled, without irony, "Beyond Human Variability" (PubMed 41760433) [WEAK — title and citation only].

So: the better the model, the more of the remaining value sits in the contested region, and the contested region is precisely what a blinded panel measures and a consensus label destroys. The four things that rise in value as the single label falls to zero are disagreement signal (where do experts diverge, and does the model diverge with them or against them), adjudicated resolution with reasoning (not the consensus answer but why the disagreement resolved that way — the training signal a majority vote deletes), hard-case and adversarial generation (cases built to break a model, which can only be authored de novo because you cannot commission a real patient with the presentation you need), and characterised uncertainty as a deliverable — the inter-rater statistics, the adjudication log, the equivocal-case policy.

The reframing

You are not selling the right answer. You are selling a calibrated map of where the right answer is knowable, where it is contested, and how far expert humans are from each other. That map does not depreciate when models get better. It is the only thing here that does not.

State the honest limit in the same breath. "Models have surpassed the expert, therefore the expert product must change" is directionally supported and not rigorously established. It is true on some narrow benchmarks and false in open-ended clinical work — the same models that beat 21 generalists on SDBench score 49.8% against pathologists' 71.8% on PathMMU (arXiv 2401.16355). Present it as a strategic hypothesis with evidence, not as a demonstrated fact.

Never sell single-annotator labels

Even when the buyer asks. Especially when the buyer asks, because they will, and it will be cheaper to produce and easier to close.

Three reasons, in ascending order of force. A single label is the commodity every one of the vendors at Centaur.ai, in full and every offshore annotation shop already sells at a third of your price. A corpus collapsed to consensus has thrown away its most valuable layer irreversibly — you cannot recover the distribution from the mean. And selling single labels trains your own buyers to price you per row, which is the market you least want to be in, because rows commoditise and evidence packages get audited.

The operational corollary: retain and version every independent assessment, blinded, with annotator identity and credentials, permanently. The disagreement is the asset. The consensus is a derived view of it.

What the first product is, concretely

Run it first in Medical coding, not in pathology. The intellectual case is strongest in pathology and the substrate is broken — only about 10% of US laboratories were digitised as of 2024 (J Pathol Transl Med), and you cannot annotate glass remotely. Coding has the same structure with none of the friction: correctness is rule-defined rather than patient-defined, so charts can be written de novo with no PHI at all; dual-coding with adjudication is already professional practice; the credential is registry-verifiable for free; and the public benchmark is wide open, with Claude Opus 5 at 63.57% as the best model in the world across 86 tested (Vals MedCode).

Note what Vals already does and does not do. MedCode's 2,755 samples were "independently annotated by two certified coders" (Vals) — two readers, no published κ, no adjudication log, no per-task variability statement. That is the gap in one sentence: the market's best public clinical benchmark is two-reader and does not report its own agreement.

The artefact. A 500-chart evaluation set: de-novo charts authored to a parameter grid, each coded independently by three credentialed coders blinded to one another, disagreements adjudicated by a fourth with recorded reasoning, shipped as data plus a reference-standard dossier — annotator credential register, grading protocol, per-task inter-rater statistics, adjudication log, equivocal-case policy. Hold back an identically constructed second set and never release it; contamination-free hold-outs are only guaranteeable with de-novo material, because every public corpus is presumptively in the pretraining data (Build it without ever touching a patient record).

What it costs. Three blinded reads plus adjudication is roughly 0.8–0.9 expert-hours per chart, plus authoring. At the observed $80–110/hr coder band that is on the order of $100–140 per fully adjudicated chart, so $50,000–$70,000 of expert labour for the 500-chart set [INFERENCE — the time assumptions are mine; the rates are observed]. No scanner, no data-use agreement, no institutional relationship. Compare the same set produced by physicians and it is five times the number, which is the entire argument for entering below the medical degree (What a clinician hour costs).

What it prices at. Two public anchors, four orders of magnitude apart, and the distance between them is the thesis.

AnchorFigureImplied unit
Tempus acquires Paige for its annotated corpus$81.25M for ~7M clinically annotated pathology slides (Fierce Biotech)$11.60 per annotated slide [INFERENCE]
Mercor builds APEX>$500,000 for 200 benchmark tasks across law, medicine, finance and consulting (Time)$2,500 per benchmark task
Stanford AIMI commercial dataset licence$70,000 per dataset per year (AIMI)a corpus, annually

Bulk annotated labels are worth about twelve dollars each in an acquisition. An adjudicated, expert-authored benchmark task is worth about two and a half thousand. Nothing about the clinician's hour explains that gap; the construction does. Price against APEX and the Stanford licence, not against per-slide annotation rates — the full pricing landscape is at What actually gets sold.

The trap

The trap is the corpus business. It looks like the same business and it is not.

A corpus is rows. Rows have a going rate, that rate is set by whoever has the cheapest credentialed labour — Loginsoft's analogue in medicine is a two-hundred-person annotation shop — and it falls every year. A corpus is also recallable: if you ever build it on data received from a covered entity, the return-or-destroy clause at 45 CFR §164.504(e)(2)(ii)(J) means your principal asset can be contractually withdrawn (Build it without ever touching a patient record).

The evidence-package business is a signed, versioned, methodologically documented artefact whose value is that an auditor can follow it. It survives commoditisation because the scarce input is not the label but the attestation, and the attestation has to be produced at collection time by a party with standing to make it.

The second trap is subtler and it is the one to watch: a buyer who wants the panel's answer without the panel's cost. They will ask for consensus labels at a single-reader price and frame it as simplification. Saying no to that is the whole strategy. The sequence for testing all of it is at Ninety days in health; the verdict it sits under is at The health read.