ClinVar holds 6,955,361 submitted records covering 4,557,411 unique variants. Of those, 165,945 carry conflicting classifications, 22,404 are expert-panel reviewed at three stars, and just 663 have practice-guideline status (statistics current 22 August 2026, NCBI). Expert panels have collectively adjudicated roughly half a percent of the unique variants in the public record.
Now put a number on the consequence. Across 1,689,845 individuals and 132,902,170 equivalent single-gene tests between September 2014 and September 2022, 41.0% of individuals had at least one variant of uncertain significance and 31.7% had only VUS results. There were 475,284 unique VUSs, of which only 7.3% were ever reclassified, at a mean time to reclassification of 30.7 months to benign and 22.4 months to pathogenic (JAMA Network Open, 25 Oct 2023).
That is a backlog of roughly 440,000 unresolved variants clearing at a decade-scale rate. It is the clearest quantified expert-judgement shortage anywhere in this dossier — and nobody has been shown to be paying to close it.
What the artefact is
Five products, and the first is the most machine-checkable expert judgement in all of clinical medicine:
- ACMG/AMP classification — a variant, the evidence codes applied (PVS1, PS3, PM2, BP4 and the rest), and the final five-tier call. Structured, auditable, gradeable. A model's output can be scored not just on the verdict but on whether it applied the right code for the right reason.
- Evidence extraction from literature — mapping a paper's functional assay to an evidence code.
- Reclassification episodes — a VUS that moved, and the evidence that moved it.
- Report generation — the clinical narrative a lab director signs out.
- Disagreement adjudication across submitting labs. ClinVar's 165,945 conflicting records are a pre-built, public, free work queue of exactly the disagreements worth adjudicating.
Product 1 is unusual in this dossier because it comes with a built-in oracle. Most clinical judgement has to be graded by another clinician; an ACMG call can be graded against a versioned, community-adjudicated public database.
Does it need patient data
No — and this is the strongest PHI position of the six, stronger even than Clinical reasoning, because the ground truth is public rather than merely absent.
A variant is not a patient. ClinVar is a US Government aggregation carrying no licence restriction, only an attribution request: "If you distribute or copy data from ClinVar, we ask that you provide attribution to ClinVar as a data source" (NCBI). The evidence base is published literature, reachable through the PMC Open Access Subset — usable per article, with the caveat that licence terms vary by article and that systematic bulk retrieval outside the sanctioned channels (PMC Cloud Service, OAI-PMH, E-utilities, BioC API) is prohibited (NLM). Same per-item filtering discipline as any image corpus.
There is no imaging, no DICOM, no IRB for public-database work, no institutional data ownership, and no licensure issue: signing out a clinical report is a CLIA lab-director act, but classification-as-training-data is not. No FDA device pathway applies in the common case. See Build it without ever touching a patient record.
ClinVar even names its own imperfection — "NIH does not independently verify the submitted information" — which is precisely a market for expert re-adjudication.
The de-novo route is barely needed here, because the public route already works. That is rare and it is worth a lot.
Is anyone buying
No. Budget scores 1, the lowest score awarded anywhere in this dossier, and the page would be dishonest at 2.
- Genomics was not evaluated at all in HealthBench Professional (PDF). Not underrepresented — absent.
- Mercor's healthcare vertical does not list genetics (Mercor). Nor does any observed rate card carry a genomics line.
- DeepMind scoped itself away deliberately. AlphaGenome beat prior SOTA on 22 of 24 sequence-prediction tasks and 24 of 26 variant-effect predictions, with 3.1–25.5% improvements — but it is API-only for non-commercial research, and DeepMind states that "our model's predictions are intended only for research use and haven't been designed or validated for direct clinical purposes" (DeepMind). The clinical-classification gap is left open on purpose.
What exists instead is a loud capability signal and one adjacent commercial relationship.
VariantBench (LatchBio) runs 118 verifiable agentic evaluations across variant calling, clinical genomics and population genetics, over 9,204 trajectories across 26 model-harness configurations. Best pass rate is 42.1% — GPT-5.6 Sol/Codex tied with Claude Opus 4.8 Max/Pi — and no model-harness pair exceeded 50% (LatchBio). A sub-50% ceiling on a well-constructed agentic benchmark is the loudest demand signal in the dossier that has not yet converted into a purchase order.
And Anthropic is already working with LatchBio: Claude for Life Sciences added Open Targets and ChEMBL connectors, and Opus 4.5 improvements were reported against LatchBio's SpatialBench (Anthropic). The organisation that built VariantBench is already inside a frontier lab's evaluation stack. That is the shortest identified path from this niche to a lab budget — see The labs as health buyers.
Academic momentum is real: AI-CURA, an automated LLM workflow for variant classification (Science Translational Medicine) [UNVERIFIED — paywalled]; GPT-4o, Llama 3.1 and Qwen 2.5 benchmarked on cancer variant classification (npj Precision Oncology); reasoning-LLM evidence extraction from clinical genomics literature (medRxiv, Feb 2026).
Proof scores 4 despite the missing buyer, because the oracle is public. An operator can curate a few thousand conflicting ClinVar variants with a named panel, publish the reclassification rate against a held-out set, and be demonstrably the best source inside a quarter — without needing a customer's permission or a hospital's data.
What the expert costs
Cost scores 4 on structure and is unscorable on labour, which is the central problem with this page.
Medscape does not report medical genetics separately [UNVERIFIED]. The relevant labour is three-tiered: ABMGG-certified clinical molecular geneticists and lab directors at physician-scale compensation; PhD variant scientists; and certified genetic counsellors, who do a large share of actual curation at a fraction of physician cost. The NSGC 2024 Professional Status Survey — which would give headcount and salary — is robots-disallowed and could not be retrieved [UNVERIFIED].
If certified genetic counsellors can produce ACMG-grade classifications at $60–90/hr against a physician's $200+, this is a business with physician-grade output at nurse-grade cost, and it is the best opportunity in the dossier. If the work in practice requires a lab director at $200/hr, it is marginal. Nothing else about this niche is as decision-relevant, and the number was not obtainable.
Everything other than labour is cheap: no imaging pipeline, no dataset licence, no de-identification, no hospital agreement. The pilot is a spreadsheet of conflicting variants, a curation protocol and a panel.
Getting to them
Reach scores 5, and unusually the best channel is not a society.
ClinGen's Variant Curation Expert Panels are the existing, named, protocolised structure for exactly this work, with a published VCEP protocol and an open application process (ClinGen). The people who do this work are already organised into blinded, adjudicating panels, publicly, by gene. And the ClinVar submitter list names every submitting laboratory, which makes both the supply side and the potential customer side directly contactable.
Add ACMG's Annual Clinical Genetics Meeting, ASHG and NSGC. ACMG's membership figure could not be retrieved from its site [UNVERIFIED], so the pool is unsized — but a channel that hands you the exact working groups by name does not need a membership count to score 5. See What a clinician hour costs.
Society posture is the friendliest in the dossier: ClinGen is a public-good consortium that wants more curation done and publishes its methodology openly. There is nothing to negotiate around.
Where the benchmarks sit
| Benchmark | Status | Frontier score |
|---|---|---|
| VariantBench (118 evals, 9,204 trajectories) | Wide open — nothing above 50% | best pass rate 42.1% |
| AlphaGenome | SOTA, non-commercial API only | 22/24 and 24/26 tasks |
| Cancer variant classification (GPT-4o / Llama / Qwen) | Open | see paper |
| AI-CURA | Claims high accuracy | [UNVERIFIED] — paywalled |
| 65 variant-effect predictors, clinical interpretation | Open | — |
Sources: VariantBench, AlphaGenome, npj Precision Oncology, Genomics 2025. Comparison at The measurement landscape.
Compare this row against Clinical reasoning, where models already beat physicians on HealthBench Professional. Here nothing clears 50%. The reference-standard ceiling problem that undermines every text niche does not apply yet.
What would kill it
Defense scores 4 and room scores 5 — nobody is in this niche, and the barriers that would keep a follower out are structural: the ACMG framework is expert-only, ClinGen membership is earned, and the backlog regenerates as fast as sequencing volume grows.
Three things kill it.
The first is that the bet never converts. VariantBench's 42% ceiling is a capability observation, not a procurement signal, and it may simply never become a line item — labs have shown no interest, and the people who care are academic consortia with no budget.
The second is concentration on the supply side. The pool is small and mostly employed by a handful of laboratories — Invitae/Labcorp, Myriad, Ambry, GeneDx and academic molecular labs — so recruiting carries genuine conflict-of-interest risk with employers who are themselves ClinVar submitters and potential customers.
The third is that the oracle cuts both ways. If ACMG classification is machine-checkable enough to grade a model on, it is machine-checkable enough for a model to be trained on cheaply from the public record — and 4.5 million public variants with 165,945 disclosed conflicts is a large free training corpus. The scarce thing has to be the adjudicated resolution, not the classification itself.
No disclosed contract exists between a frontier lab and a specialist-physician data vendor anywhere on the diagnostic side, and in genomics there is not even inferential evidence to extrapolate from — no rate card, no benchmark acknowledgement, no job posting. The buying case rests entirely on a failing benchmark and one adjacent vendor relationship.
Where the record is thin
Three holes, in order of consequence. The genetic-counsellor labour cost is unknown and decides the whole niche. ACMG membership is unknown, so the pool is unsized. And AI-CURA's accuracy claims are behind a paywall, which matters because if an automated workflow already classifies variants at high accuracy, the human curation product shrinks to adjudication of the residual.
Ranked comparison at The health read; the strategic read at Characterised disagreement; the market backdrop at Expert data for frontier labs.