Eight board-certified dermatopathologists with a median 34 years' experience reviewed 792 melanocytic slides from 736 patients across eight German university hospitals. They agreed unanimously in 53.5% of cases. In 9.1% there was no majority at all. Overall Fleiss' κ was 0.701 — but only κ = 0.428 for non-invasive melanomas, and the expert panel disagreed with the original local pathologist in 14.9% of cases (Nature Communications, 17 Jan 2025).
That is not a quality problem to be fixed. It is what the ground truth is.
Once you accept it, the product follows without argument. When truth is a distribution over expert opinion rather than a fact, the sellable thing is not a label — it is a calibrated panel, with known reader identities, documented experience levels, blinding between readers, and a recorded adjudication of the disagreements. That is a service, it repeats, it cannot be crowdsourced, and it maps almost line by line onto what the FDA started asking for in January 2025.
The problem is underneath all of it: only about 10% of US laboratories were digitised as of 2024 (J Pathol Transl Med). You cannot annotate a glass slide remotely.
What the artefact is
Five products, in ascending order of defensibility:
- Slide-level and region-level diagnosis labels on whole-slide images — the commodity end.
- Cell, nucleus and gland segmentation masks.
- IHC scoring — HER2 0/1+/2+/3+, PD-L1 CPS/TPS, Ki-67 index. Ordinal, protocol-defined, and known to be unreliable between humans, which makes it the most defensible consensus-panel product in all of medicine. PD-L1 CPS in upper-GI adenocarcinoma is documented as having "high interobserver variability among pathologists" (Modern Pathology) though the exact κ was not retrievable
[UNVERIFIED]; HER2-low IHC has documented interobserver and inter-antibody irreproducibility (PubMed 36856777)[UNVERIFIED]on figures. - Multi-reader consensus panels with adjudication — N pathologists score independently, disagreements are adjudicated, and the distribution itself is the label.
- Synoptic reporting against CAP cancer protocols, as structured text.
Product 4 is the business. Products 1 and 2 are what Centaur Labs already sells — see Centaur.ai, in full.
Does it need patient data
Less than you would expect, and the binding constraint is not privacy but supply.
TCGA whole-slide images are distributed through TCIA, where most collections carry CC-BY licences permitting commercial use; a small number restrict it, and brain and head/neck sets sit under a TCIA Limited Access License (TCIA). CAMELYON is public. That is a real, licence-clean commercial substrate — but note the two-tier structure at the NCI Genomic Data Commons: the open tier is summary-level and exhaustively mined, while the controlled tier requires dbGaP authorisation and Data Access Committee approval keyed to a stated research purpose, which fits a commercial annotation vendor badly. NCI is also explicit that "TCGA has no rights to redistribute materials outside of the program" — the physical tissue is gone.
Look at how MedGemma 1.5 was actually built: trained on TCGA and CAMELYON plus four private de-identified H&E whole-slide datasets from a European academic hospital, a US commercial biobank, a CRO and a tertiary teaching hospital (model card). Two public sources, four bought. That ratio is the market.
The de-novo route exists here but is weaker than elsewhere: a pathologist can write a case narrative or adjudicate a described morphology, but the value of pathology judgement is bound to the pixels. Slides and their digital surrogates are institutional property — in every US state except New Hampshire the provider, not the patient, owns the medical record — so a pathologist cannot bring their own material. The route that avoids PHI entirely is panel adjudication over already-public slides, which is legitimate, cheap and licence-clean. See Build it without ever touching a patient record; note its wider conclusion that the corpora which are commercially open are open to everyone and therefore carry no moat.
CLIA adds a deployment-side constraint that shapes who buys rather than what you may hold: CAP notes AI device performance must be verified locally and calibration verified every six months (CAP response to HHS AI RFI, 23 Feb 2026).
Is anyone buying
Budget scores 2, and the honesty of that number is the point.
Frontier labs: almost nil. Pathology contributed 3 of 525 examples to HealthBench Professional (PDF). Anthropic ships an Owkin pathology connector (Anthropic) — routing to a pathology tool rather than building pathology capability. See The labs as health buyers.
Foundation-model builders buy pixels, not judgement. Virchow, UNI, Prov-GigaPath and peers are largely self-supervised on unlabelled slides; the entire design goal is not needing pathologist labels at scale (Nature Communications). Pathologist time is needed for evaluation and downstream task labels, not pretraining — a smaller, later, more episodic budget.
Device sponsors do buy. FDA-cleared or authorised pathology AI now includes AISight Dx (PathAI), PathPresenter Clinical Viewer, Roche Digital Pathology Dx, Paige Prostate, Ibex Prostate Detect, Artera AI, RlapsRisk BC and CHAI (J Pathol Transl Med). Each needed an adjudicated pathologist reference standard, and each future one will need it documented to the January 2025 standard — The regulator wrote your product spec.
The one disclosed price in the sector is an acquisition, not a contract. Tempus acquired Paige for $81.25M in August 2025, explicitly for its "proprietary dataset of almost 7 million digitized pathology slides that are clinically annotated" (Tempus IR). That is the clearest valuation of annotated pathology data on record — and it values it as a company to be bought, not a service to be retained.
Observed labour rate: $100–200/hr for a licensed pathologist on a whole-slide super-resolution evaluation task (OpenTrain).
Proof scores 3. A κ-and-adjudication study over public TCGA slides is publishable and would establish authority quickly — the melanocytic paper is the template. But it takes a study cycle, and the audience for it is device sponsors rather than labs, which is a slower procurement.
What the expert costs
Cost scores 3. Pathologists average $394K, implying $197–219/hr [INFERENCE] — the lowest opportunity cost of any physician specialty here, and therefore the best wage-arbitrage position in the dossier (Becker's/Medscape). Observed AI rates of $100–200/hr sit just under that, which is thin but workable for marginal hours — see What a clinician hour costs.
What pushes cost down from 4 is the panel itself. A defensible product needs N readers per case plus an adjudicator, so the labour multiplies by three to eight before any margin. Market comparators: MD Anderson charges a minimum $571 per consultation on referred slides at 48-hour turnaround (MD Anderson), and annotation vendors quote $200–800 per whole slide (mpiricsoftware) [WEAK — single vendor blog].
Getting to them
Reach scores 4. CAP exceeded 20,000 members for the first time in 2025 and accredits more than 8,400 laboratories across 60 countries, including over 750 outside the US and Canada (CAP Annual Report 2025). An NPPES-derived count puts 24,869 actively practising US pathologists (hyperdrivebio) [WEAK — vendor, derived].
The accredited-lab list doubles as a B2B channel: it tells you which labs are digitised, which is the scarce variable. USCAP and CAP annual meetings reach the individuals. And CAP's proficiency-testing programme is already the national infrastructure for scoring pathologist agreement — simultaneously the most valuable partnership target and the most credible competitive threat.
Society posture is constructive rather than hostile. CAP asks for "manufacturer transparency regarding AI tools," flags "potentially imbalanced datasets for training" and insists pathologists remain the physician leaders of the laboratory. A well-run panel product is aligned with CAP's stated position, which is unusual and worth a great deal.
Where the benchmarks sit
| Benchmark | Status | Scores |
|---|---|---|
| PathMMU | Open, large human–AI gap | 33,428 MCQs, 24,067 images, validated by 7 pathologists; GPT-4V zero-shot 49.8% vs human pathologists 71.8% |
| PathView-Bench | New, open | [UNVERIFIED] |
| Public pathology foundation models, clinical benchmark | Open | — |
| MedGemma 1.5 pathology | — | per-task scores not headlined |
Sources: PathMMU, PathView-Bench, Nature Communications. Context at The measurement landscape.
PathMMU is the one benchmark in this dossier where humans clearly beat the model — 71.8% against 49.8% — which is the opposite of the clinical-reasoning picture and a reason the ground-truth premise survives longer here than at Clinical reasoning.
What would kill it
Defense scores 5 — the highest in the dossier — because none of the usual erosions apply. Consensus panels cannot be crowdsourced; the FDA's annotator-expertise and adjudication requirements do not expire; IHC assays keep changing, so the panels need re-running; and a model that beats one pathologist still cannot be the distribution over eight. Room scores 4: Paige has been absorbed into Tempus, Centaur Labs sits at the commodity end, and nobody owns adjudicated-panel-as-a-service.
What kills it is the substrate. Ten percent digitisation is not something an operator can fix, and it caps addressable volume regardless of how good the panel is. Three of 525 HealthBench Professional examples says the largest buyer is not here yet. And if pathology foundation models keep improving self-supervised, the label budget stays small and episodic.
No contract between a frontier lab and a specialist-physician data vendor has been disclosed anywhere on the diagnostic side. The Paige acquisition price is a valuation of an annotated corpus, not of a labelling service. All buying evidence on this page — device-sponsor need, Owkin routing, the OpenTrain rate — is inferred from rate cards, clearance lists and connector announcements.
Where the record is thin
The IHC variability figures that carry the most weight in the thesis — PD-L1 CPS and HER2-low κ values — could not be retrieved from their primary sources. Only the melanocytic study gives real numbers. USCAP attendance is not published [UNVERIFIED], so channel sizing is soft. The $200–800 per-slide annotation price rests on a single August 2026 vendor blog and should not carry weight in a pricing decision without a second independent quote.
Ranked comparison at The health read; the wedge argument at Characterised disagreement; the generalist backdrop at Expert data for frontier labs.