Miju Labs

All niches

Clinical medicine

The best-documented lab demand in the set and the widest open benchmark among the professional domains — attached to the highest cost basis anywhere, where a gastroenterologist's opportunity cost exceeds Mercor's entire ceiling.

watchhigh confidence7 minupdated 2026-08-30
Who the expert is
Board-certified physicians; primary care, paediatrics and psychiatry are the affordable tiers
What they earn by day
$132.66/hr median (BLS, physicians and surgeons)
What data work pays
Locum tenens average $215/hr — the true opportunity cost of a physician's marginal hour
Size of the pool
862,800 US physicians
How you reach them
ABMS diplomate registers, state medical board lookups, Doximity, Sermo, Figure 1, Physician Side Gigs
Benchmark position
MedQA saturated above 95%; HealthBench and MedHELM wide open (MedHELM leader 0.652)
Who is already there
Centaur AI, Protege (labelled artefacts only — reasoning traces are unoccupied)
Read
Everything about the demand says build. The wage floor says you can only afford the specialties the buyer cares least about — solve that and this is the best niche here.
Speed to proof
4
Budget now
5
Defensibility
4
Cheap to start
2
Reachability
5
Room to win
4

Medicine is the only domain in this atlas where a frontier lab has published, in detail, exactly how it buys expert judgement — and then priced itself out of your reach on the same page.

OpenAI built HealthBench with 262 physicians who have practised in 60 countries, spanning 49 languages and 26 specialties, producing 5,000 realistic health conversations and 48,562 unique physician-written rubric criteria (OpenAI). Google DeepMind's AMIE appeared in Nature: a randomised, double-blind crossover OSCE with 20 board-certified primary care physicians, 20 validated patient-actors and 159 case scenarios across Canada, the UK and India, graded by majority vote of three specialists across 32 axes — AMIE beat the physicians on 30 of 32 specialist-rated axes (Nature), with multimodal follow-on work in Nature Medicine (2026).

That is a complete, replicable specification of the product, published by the buyer, twice.

Then the cost. Physicians' median is $132.66/hr across 862,800 US jobs (BLS), but the number that binds is locum tenens, which is what a physician's genuinely marginal hour clears at: $215/hour on average, with gastroenterology at $367, anaesthesiology $292, radiology $289, cardiology $272, emergency medicine $258, psychiatry $223, family medicine $140 and general paediatrics $108, across a full range of $60–$500/hr (Physician Side Gigs, June 2025). Mercor's advertised expert ceiling is $200+/hr. A gastroenterologist's alternative use of the hour exceeds the ceiling of the largest buyer in the market.

What the data actually is

Every existing medical-data vendor sells labelled artefacts — images, scans, records. None sells physician reasoning traces, diagnostic deliberation or specialist second opinions as structured expert data. That is the open position, and it is the expensive one.

Concretely:

  • Rubric criteria, HealthBench-style. 48,562 of them written by physicians is a purchasing specification, not a research artefact — it tells you what OpenAI paid for and how much of it they needed.
  • Diagnostic deliberation. The differential, the order in which it was considered, the finding that collapsed it. Not the answer.
  • Second opinions with disagreement preserved. Where two specialists diverge on the same vignette, and why. AMIE's three-specialist majority vote implies exactly this data underneath it.
  • Phase decomposition. Differential → workup → management, scored separately. The phase-decomposition move: a leaderboard tells a lab it is third, which is unwelcome; a phase breakdown tells it where the model collapses, which is a purchase order.

None of it requires PHI, which is the whole trick — see below. Proof scores 4: the artefact is unoccupied and MedHELM gives you a public number to beat, but HealthBench's 262-physician scale sets a bar a small cohort cannot clear quickly.

Is anyone buying

Yes, more explicitly than anywhere except cyber. Budget scores 5.

  • OpenAI is hiring Research Engineer/Scientist, Health AI at $310,000–$460,000, Software Engineer, Healthcare at $245,000–$385,000, plus Forward Deployed Engineers (Healthcare) in NYC, SF and Seattle (Becker's, OpenAI careers). Health and cyber are the only two domains in this atlas with named data-producing reqs carrying pay bands.
  • xAI named medicine as one of four specialist tutor domains — "STEM, finance, medicine, safety" — when it surged its specialist team 10x (TechCrunch).
  • Mercor's APEX used Mount Sinai Hospital graders (TIME).
  • Vals AI benchmarks medicine alongside law, banking and engineering, and its results appear in model cards from OpenAI, Anthropic, Google, Meta and xAI (Pulse 2).

The benchmark side is equally clear. MedQA is saturated — frontier models above 95%, up the chain from PaLM's 67.6% through Med-PaLM 2's 86.5% (Kili survey). HealthBench is open: GPT-3.5 Turbo 16% → GPT-4o 32% → o3 60%, with OpenAI saying HealthBench Hard leaves "plenty of headroom" (OpenAI). MedHELM (Stanford CRFM and Pacific AI, 40+ clinical scenarios, Q2 2026) has Gemini 3.1 Pro leading at 0.652, then Gemini 3.5 Flash 0.642, Muse Spark 0.621, GPT-5.4 mini 0.552 and GPT-5.4 0.538 (Pacific AI).

Proven budget and an open benchmark is the combination The specialist wedge says you almost never get — cyber and code have budgets with captured benchmarks, law and accounting have empty benchmarks with no proven budget. Clinical breaks the pattern. Cost is the price of that exception.

Getting the experts

Reach scores 5. Medicine is the most institutionally organised profession in this set: ABMS specialty board diplomate registers are a verifiable credential, state medical board licence lookups confirm standing, and Doximity has near-universal US physician adoption.

Better still are the communities built for exactly the behaviour you want to elicit. Sermo and Figure 1 are physician-only case-discussion platforms — structurally identical to your product, already populated. And Physician Side Gigs, a ~100k+ physician moonlighting community whose entire premise is finding non-clinical paid work, is the single best-fitted channel named anywhere in this atlas: an audience pre-sorted for willingness to sell marginal hours.

Then locum agencies (Weatherby, All Star, CompHealth), specialty society meetings (RSNA, ACEP, ACC), residency and fellowship programme directors, and r/medicine and r/Residency.

What it costs to run

Cost scores 2, and this is where the niche is decided.

You must segment by specialty, because the opportunity-cost curve is steeper than in any other domain. General paediatrics at $108/hr and family medicine at $140/hr are affordable and abundant; psychiatry at $223 is borderline; radiology at $289 and gastroenterology at $367 are not payable at all on an hourly basis against any observed AI-data rate.

The workaround is to stop buying hours. For procedural specialties, pay per artefact for high-value second opinions — a structured disagreement on a hard vignette is worth a flat fee a specialist will accept for twenty minutes of work, where the same person would refuse $200/hr for four hours. That is the bounty logic imported into medicine, and it is the only version of this business that clears.

Compare Accounting, audit and tax at $40.23/hr for a credentialed, register-listed professional. On the same $150/hr sell price, the CPA business has a spread and the gastroenterology business has a loss. See What a rake can actually be.

Who is already there

Room scores 4 because the incumbents sell a different thing. Centaur Labs (now Centaur AI) raised a $15M Series A led by Matrix and then a $16M Series B for crowdsourced medical annotation via its DiagnosUs app (Centaur, BioSpace) — see Centaur Labs. Protege raised $30M led by a16z for a governed training-data marketplace with a healthcare vertical and imaging partnerships with Segmed and Gradient Health (SiliconANGLE). OmicsBank raised $2.25M for clinical data infrastructure explicitly naming frontier AI labs (PR Newswire).

Hippocratic AI, at a $126M Series C and a $3.5B valuation with a clinician safety-supervision network (Business Wire), is a product company — a customer for your data, not a competitor.

Defense scores 4. The credential takes a decade to obtain and cannot be faked, the compliance overhead deters casual entrants, and clinical guidelines revise continuously — the supply stays scarce and the data keeps needing refreshing.

What would kill it

What would kill it

The wage floor, first and most likely. If the buyer wants radiology and cardiology judgement, you cannot assemble it at a price that leaves a spread, and you will discover this in month five with a cohort of paediatricians and a customer who wanted oncologists.

Malpractice. A physician who writes a diagnostic reasoning trace that later trains a model implicated in patient harm carries a novel, unpriced liability. Carriers have no product for it. One widely shared story about that exposure removes your supply.

PHI, if you ever touch it. Real charts, images and notes cannot leave a covered entity without a BAA, Safe Harbor or Expert Determination de-identification, and often IRB review.

Licensure and employer claims. Judgement is licensed per state, which makes a national product jurisdictionally messy, and health systems increasingly assert rights over clinical data and, in some contracts, physician work product.

The escape is visible in the evidence: HealthBench used synthetically generated, adversarially tested conversations with physician-written rubrics, and AMIE used patient-actors, not patients. Neither touched PHI. Build synthetic vignettes plus physician rubrics and the HIPAA problem evaporates entirely — leaving only the cost problem, which is the real one.

The first ninety days here

Recruit through Physician Side Gigs and Sermo, not LinkedIn. Verify against ABMS diplomate registers on day one. Start with primary care, paediatrics and psychiatry — the affordable band — and build 300 synthetic vignettes with three-physician rubric grading, MedHELM-comparable so the number means something to a reader.

Then run one per-artefact pilot with a procedural specialty to test the flat-fee hypothesis before you build anything on it. If specialists will not take a bounty, this niche is a primary-care business and should be priced as one. See The first ninety days.

Where the record is thin

Locum rates are a June 2025 snapshot from a physician community's own survey, and they vary enormously by geography and urgency.

No lab has published what it paid HealthBench's 262 physicians, so the sell-side price for rubric criteria is unknown and every margin estimate here is inferred. Whether physicians will accept per-artefact payment at all is untested — the entire cost workaround rests on an analogy to bug bounty, in a profession with the opposite risk culture.