Miju Labs

The security dossier

Oncology

The largest and most internationally reachable clinician pool of the six, a severe documented workforce shortage — and every road to the valuable artefact runs through a closed door.

avoidmedium confidence8 minupdated 2026-08-30

Oncology should be the obvious niche. The supply is the largest and most internationally distributed of the six — ASCO has more than 50,000 members across more than 150 countries, roughly a third of them outside the US (ASCO). The clinical need is severe and worsening: oncologist density fell from 15.9 to 14.9 per 100,000 people aged 55+ between 2014 and 2024, 68% of the 55+ population lives in counties with at-risk oncologist coverage, 11% of older Americans live in rural "cancer care deserts," and by 2037 non-metropolitan areas are projected to meet only 29% of demand against 102% in metro areas (ASCO, 7 October 2025).

And then the artefact that carries the value turns out to be someone else's copyright.

NCCN Guidelines are closed. Permission is required for any reproduction or distribution; there is no published commercial, derivative-work or machine-readable licence; the only route is an email to permissionrequest@nccn.org (NCCN). An oncology expert-data product whose value proposition is "NCCN-concordant" is building on land it does not own.

What the artefact is

Six products, and the ranking of their value is almost exactly the inverse of their availability:

  1. Staging — AJCC TNM assignment from a written case. Rule-following, gradeable, cheap, and producible entirely de novo.
  2. Treatment planning — line of therapy, regimen, dose, sequencing.
  3. Guideline application — pathway traversal with the decision node cited. The most valuable, and the one that requires the closed guidelines.
  4. Molecular tumour board reasoning — variant to actionability to trial matching.
  5. Multidisciplinary tumour board transcripts and consensus recommendations — the closest thing in medicine to a naturally occurring multi-agent episode, and therefore the most interesting artefact in the entire dossier for anyone training agents. Also the most institutionally locked.
  6. Trial eligibility adjudication.

Items 1 and 6 are clean. Item 3 is legally fraught. Item 5 is effectively unobtainable without hospital BD.

Does it need patient data

Mixed, and awkwardly so — this is the least clean PHI position of the six.

Staging and guideline application can be produced from de-novo written cases, which touches no PHI, engages no covered entity, creates no human subject and creates no institutional data-ownership claim. That route works, and it is cheap. See Build it without ever touching a patient record.

But real multidisciplinary tumour board material — the artefact that would actually differentiate an oncology product — is PHI-dense and institutionally owned. A tumour board transcript names a patient, a disease, a treatment history and a date. Safe Harbor de-identification handles the direct identifiers but item (R), "any other unique identifying number, characteristic, or code," is a standard rather than a list, and it swallows most of what makes an oncology narrative valuable: a rare histology, an unusual treatment sequence, a distinctive molecular profile. The date rule is worse. Only the year survives, and an oncology case whose value is in the interval — time to progression, sequencing of lines, response duration — is gutted.

Then the licensing blocker, which exists nowhere else in this dossier. Even a perfectly PHI-free de-novo case becomes commercially fraught the moment its rationale is expressed as NCCN pathway traversal. TCGA is available for the molecular side, but the open tier is summary-level and already exhaustively mined, and the controlled tier runs through a Data Access Committee keyed to a stated research purpose.

Practically: you can build a clean staging and eligibility product with no patient data at all, and you cannot build the guideline product at all without a negotiated licence.

Is anyone buying

Budget scores 2. There is evidence, and all of it is one step removed.

  • HealthBench Professional lists haematology/oncology among its top-represented specialties — the strongest direct evidence that a lab has paid oncologists (PDF). But it is a specialty-mix bullet, not a contract, and the artefact is general clinical text rather than oncology-specific reasoning.
  • Mercor recruits oncology and haematology explicitly in its healthcare vertical (Mercor) — which means the generalist marketplaces already cover this supply.
  • OpenAI markets an oncology use case rather than buying oncology data. Color Health built a cancer copilot on GPT-4o with clinical input from Stanford and UCSF oncologists and the American Cancer Society CEO; clinicians using it identified 4× more missing labs, imaging or biopsies and cut analysis time to 5 minutes from weeks (OpenAI). That is a customer story.
  • Anthropic reaches oncology through partners — Flatiron Health is a named Claude for Healthcare partner (Anthropic). See The labs as health buyers.
  • There is a healthy benchmarking literature with no lab co-authors. GPT-4 for molecular tumour board recommendations (The Oncologist); retrieval-augmented GPT-4 improving NCCN-concordant lymphoma recommendations (Blood/ASH); CONCORDIA, a blinded prospective study benchmarking LLMs against urological MDTs (ScienceDirect). Academic oncologists are producing this evaluation work, which means the labs are getting it free.

Proof scores 2, the lowest in the dossier. There is no leaderboard to top and no canonical eval to beat, so there is no fast public route to being visibly the best source. You would have to build the benchmark first — and the benchmark everyone would want is the NCCN-concordance eval, which is exactly the one the copyright forecloses. That is presumably why nobody has built it.

What the expert costs

Cost scores 2. Oncology and haematology average $464,000, implying $232–258/hr of opportunity cost [INFERENCE] (Becker's/Medscape). No oncology-specific AI-data rate has been observed anywhere; Mercor's disease-area clinician band of $150–230/hr is the only usable proxy, and it sits below opportunity cost.

The workforce data makes this worse, not better. A specialty with 68% of its elderly catchment in at-risk coverage counties is a specialty whose members have no spare hours. The counterweight is genuine: community oncologists in under-served areas are simultaneously the most time-poor and the most motivated by supplemental remote income, and they are the recruitable segment. But you are buying evening hours from people who are already overloaded, at a rate below what their clinical hour is worth. See What a clinician hour costs.

Add legal cost. Treatment recommendations are the highest-stakes output in this dossier, so the contributor agreement and the framing of every artefact as non-clinical needs more work here than anywhere else — What the doctor on the other end is risking.

Getting to them

Reach scores 5, the one unambiguously excellent number on this page.

ASCO's 50,000+ members across 150+ countries is the largest and most internationally reachable pool of the six, and the international third matters more here than anywhere else because it decouples recruiting from US compensation levels. Add ASH for haematology, ESMO for Europe, the NCCN member institution list, and the community-oncology segment described above.

There is a real strategic argument in that geography: an operator that recruits internationally is buying oncology judgement at a fraction of the US $232–258/hr opportunity cost while still producing English-language, guideline-anchored artefacts. It is the only niche in this dossier where the supply geography is itself the business model.

Where the benchmarks sit

This is the weakest benchmark landscape of the six, and the weakness is structural rather than incidental. There is no canonical public oncology benchmark analogous to ReXrank, PathMMU, HealthBench or VariantBench. What exists is a scatter of single-institution retrospective concordance studies:

StudyDomainDesign
The Oncologist 2025Molecular tumour board recommendationsGPT-4.0 against real-world data
Scientific Reports 2026Anaplastic thyroid cancerguideline adherence across leading LLMs
PubMed 42462288Gynaecologic oncologyguideline-anchored RAG pre-integration benchmark
PubMed 42346209Colorectal MDTgeneral-purpose vs domain-specialised models
CONCORDIAUrological oncologyblinded, prospective, against real MDTs

No leaderboard, no shared test set, no version control, no public scores to point at. Compare The measurement landscape for what a contested benchmark does for a niche — and note that its absence here means an operator cannot demonstrate superiority, cannot be cited in a system card, and cannot use benchmark authorship as marketing.

The absence of a canonical oncology benchmark is itself the opportunity, and it is a trap. Nobody owns "the NCCN-concordance eval," and building it would be a credible wedge — except that the NCCN licensing position makes it legally fraught. The gap is not an oversight. It is a consequence.

What would kill it

Defense scores 3 and room scores 4. Room is high because nobody is there; it is not 5 because the reason nobody is there is a live legal blocker, not an oversight anyone can out-execute.

Four kill mechanisms, in order of severity.

NCCN copyright is the dominant, specific and under-appreciated one. Every high-value artefact routes through it, and the counterparty is a consortium with no published commercial licence and no evident commercial incentive to grant one.

Malpractice exposure is the highest of any output in this dossier. A rendered oncology treatment opinion on an identified patient is materially more exposed than a radiology read, and the "curbside consult" doctrine is eroding — courts have found duty in informal consultations (ASCO Post). Annotation of de-identified or de-novo material creates no physician-patient relationship, but the boundary has to hold under harder facts here than anywhere else.

No regulatory pull. Oncology decision support of this kind mostly sits inside the Clinical Decision Support carve-out, so it does not attract the FDA's documented-annotator-expertise requirement the way an imaging device does. Radiology gets structural demand from The regulator wrote your product spec; oncology does not.

Guideline churn cuts both ways. Treatment standards change constantly, which means the data needs refreshing — genuinely good for defensibility, and the reason defense is 3 rather than 2 — but each refresh runs back through the same closed door.

No disclosed contract exists anywhere

No disclosed contract exists between any frontier lab and any specialist-physician data vendor on the diagnostic side of medicine, and oncology has the thinnest inferential evidence of the six. "Haematology/oncology is a top-represented specialty in HealthBench Professional" is a line in a specialty-mix chart. There is no dollar value, no term, no named vendor, and no oncology-specific rate card anywhere in the record.

Where the record is thin

There is no observed oncology AI-data rate at all, so the entire cost model rests on a generic disease-area proxy. It is unknown whether NCCN has ever granted a machine-readable or derivative-work licence to anyone, and that single answer would change the ranking of this page materially — it is worth an email before anything else. ASCO annual-meeting attendance and the size of the recruitable community-oncology segment are both unsized.

Ranked comparison at The health read; the strategic read at Characterised disagreement; the credentialled-professions backdrop at Clinical medicine and Expert data for frontier labs. For what a specialist incumbent looks like once one exists, see Centaur.ai, in full.