Miju Labs

The security dossier

The measurement landscape

Models now beat physicians on HealthBench Professional, 59.0 against 43.7, and by four times on SDBench — while humans still beat GPT-4V on pathology. When the model exceeds the expert, the buyer stops wanting ground truth and starts wanting harder cases.

high confidence9 minupdated 2026-08-30benchmarks · healthbench · medhelm · vals · evaluation · saturation

Benchmarks are the sales artefact in this market. A lab that is publicly bad at something has a budget for it; a benchmark that is saturated has no buyer. So the map of what is measured, and how badly models do on it, is the map of where the money can go.

The state of every health benchmark that matters

BenchmarkOwnerStateBest score
HealthBench ConsensusOpenAINear-saturatedgpt-5-thinking-mini 96.5%
HealthBench Hard (1,000 examples)OpenAIFar from saturatedgpt-5-thinking 46.2%; o3 31.6%; GPT-4o 0.0%
HealthBench Professional (525 examples)OpenAIOpen, newGPT-5.4 in ChatGPT for Clinicians 59.0; base 48.1; human physicians 43.7
MedHELM — Clinical Decision SupportStanfordWeakest category0.56–0.72 across nine frontier models
MedHELM — Administration & WorkflowStanfordWeak0.53–0.63
MedHELM — Note GenerationStanfordStronger0.73–0.85
Vals MedScribeVals AI + ProtegeSaturatingClaude Opus 5 90.99%, leaders within ~3 points
Vals MedCode (ICD-10-CM)Vals AI + ProtegeWide openClaude Opus 5 63.57%; Gemini 3.1 Pro 59.06%
VariantBenchLatchBioWide openbest pass rate 42.1%; nothing above 50%
BELO (ophthalmology)AcademicSplitaccuracy 0.882; reasoning quality 20.40–71.80/100
SDBench (304 NEJM CPC cases)Microsoft AIOpenMAI-DxO 80% (85.5% best config) vs 21 physicians at 20%
PathMMUAcademicHumans still aheadGPT-4V 49.8% vs pathologists 71.8%
MedXpertQAAcademicContestedleaderboard top 0.784
NursingEssentially emptyone synthetic MCQ set
mPACT (behavioural health)mpathicProprietaryClaude Sonnet 4.5 led on clinical alignment

Sources: GPT-5 system card, HealthBench Professional, MedHELM, Vals MedScribe, Vals MedCode, VariantBench, BELO, SDBench, PathMMU, NurseLLM.

Four readings come off that table before anything else.

HealthBench is not one benchmark, it is three, and only one is alive. Consensus is done at 96.5%. Hard sits at 46.2% and GPT-4o scored a literal zero on it. A benchmark family with a saturated tier and an unsaturated tier is the ideal shape for a vendor, because the saturated tier proves the instrument works and the unsaturated tier justifies the next contract.

The weakest MedHELM category is the one that matters clinically. Clinical Decision Support at 0.56–0.72 across nine frontier models, against Note Generation at 0.73–0.85. MedHELM's taxonomy — 5 categories, 22 subcategories, 121 tasks, validated by 29 clinicians across 14 specialties — is the structural map of the field, and the open capability gap sits exactly where clinical judgement lives (MedHELM, Nature Medicine).

The Vals pair is the cleanest natural experiment in the sector. Same builder, same data licensor, same construction method, two tasks: note generation at ~91% and diagnosis coding at ~64%. A 27-point gap that has persisted across snapshots — the earlier figures were 88% and 56% — which means the gap is the durable fact, not the levels. Documentation is nearly solved; coding is not. See Medical coding and Ambient scribing.

Two spaces are effectively empty. Nursing has no established benchmark at all: MedHELM has no nursing category, and the only nursing-specific artefact is NurseLLM, which built its MCQ set through "a multi-stage data generation pipeline" — synthetic, not nurse-authored (arXiv). Genomics fares no better on the buyer side: VariantBench's frontier ceiling of 42.1% with nothing above 50% is a loud capability signal attached to no observed purchaser. Nursing and Genomics and variant curation take those apart.

BELO is the costable recipe

One benchmark itemises its own expert labour, and that makes it the most useful single artefact in the measurement landscape.

BELO — 900 ophthalmology MCQs drawn from BCSC, MedMCQA, MedQA, PubMedQA and BioASQ — used roughly 23 ophthalmology professionals across three tiers: one board-certified ophthalmologist plus two optometrists plus six research staff for quality control; ten board-certified ophthalmologists refining explanations; three senior ophthalmologists adjudicating (arXiv, Ophthalmology Science).

That is a repeatable, costable production recipe — the only one published at this level of detail. And what it produced is the split that reframes the product:

The BELO split

OpenAI o1: accuracy 0.882, macro-F1 0.890. Text-generation quality: 20.40–71.80 out of 100.

The answer is nearly right and the reasoning is nearly worthless. Accuracy is approaching its ceiling while reasoning quality is wide open. Sell reasoning-quality labels, not answer-correctness labels — the correctness market is closing and the reasoning market has barely opened.

The models have passed the physicians

This is the finding that changes what the product should be.

On HealthBench Professional, ChatGPT for Clinicians running GPT-5.4 scores 59.0 against human physicians' 43.7 (OpenAI). The base model at 48.1 also beats the physicians. On SDBench, Microsoft's MAI-DxO reached 80% (85.5% in its best configuration) against 20% for 21 practising physicians — a four-fold gap — at 20% lower cost than the physicians and 70% lower than standalone o3 (arXiv).

The counter-example is real but narrow. On PathMMU — 33,428 multimodal questions over 24,067 images, validated by seven pathologists — GPT-4V scored 49.8% zero-shot against 71.8% for human pathologists (arXiv). Where the task is genuinely multimodal and specialist, humans are still comfortably ahead. Where it is text-mediated clinical reasoning, they are not.

Two caveats keep this honest. The physician baselines in these studies are constructed under artificial conditions: SDBench's 21 generalists worked without colleagues, references or the usual escape hatches, and MAI-DxO's headline comparison rests on about 149 physician-hours in total. And a benchmark on which the model beats the expert may be measuring the benchmark's own construction as much as the model's competence — HealthBench Professional's rubrics were written by the same physician population being outscored.

But directionally the trend is not ambiguous, and it survives the caveats.

What that does to the product

The reframing

When the model exceeds the expert, the buyer stops wanting ground truth and starts wanting harder cases and disagreement signal. Ground truth is a commodity the model can now approximate. What it cannot generate is the case it gets wrong, or the reasoned disagreement between two specialists who both know what they are doing.

Every part of the market already shows this shift in its construction, even where nobody has named it.

HealthBench Professional selected 525 benchmark examples out of 15,079 candidates, with difficult cases enriched roughly 3.5× — the value is in the filtering, not the volume. HealthBench Hard exists as a separate 1,000-example tier for exactly this reason. BELO's three-tier adjudication with senior tiebreak is a disagreement-harvesting machine sold as a quality-control process. MedCode's two independent certified coders per sample produces, as a by-product, a record of every case two professionals coded differently — and that by-product is more valuable to a post-training team than the consensus label it was built to produce.

That is a different product from ground truth, and a better one. It is more expensive per unit — two independent reads plus adjudication rather than one pass — so it carries a higher price naturally. It resists commoditisation, because the cases a model fails are defined relative to a model and change with every release, which makes it a subscription rather than a one-off corpus. It is exactly what the FDA asks AI-device sponsors to document by name (The regulator wrote your product spec). And it is the one thing a lab cannot self-supply cheaply, because generating hard cases requires knowing where the model breaks — which is the buyer's own information, and the only thing they will trade for it is money.

The counter-argument deserves stating. If models keep improving at this rate, the population of cases where a specialist adds signal shrinks year by year, and the business is selling into a narrowing band. That is the What better models do to each layer question in its sharpest health-specific form, and nothing in the record settles it. What the record does show is that the narrowing is uneven: MedCode at 63.57%, VariantBench at 42.1%, MedHELM Clinical Decision Support at 0.56, PathMMU with a 22-point human lead. There is a great deal of unnarrowed band left.

The unmodelled problem

Nobody has published evidence that buyers have made this shift. The reference-standard ceiling is visible in the scores, and no lab has said in writing that it now wants disagreement signal rather than gold answers. Whether the reframing is a year early or a year late is the single most consequential unknown in the product section — What the health dossier could not establish.

The read

Saturation is uneven and mapped: HealthBench Hard at 46.2%, MedCode at 63.57%, VariantBench at 42.1%, MedHELM Clinical Decision Support at 0.56, nursing at nothing. Those are the addressable gaps. But the deeper shift is that physician-generated ground truth has stopped being the ceiling on text-mediated reasoning, and the product that replaces it — harder cases, adjudicated disagreement, reasoning-quality labels — is more defensible than the one it replaces. That is the argument Characterised disagreement is built on.