Miju Labs

The security dossier

What the health dossier could not establish

Ranked by how much the answer moves the decision. The top item is that nothing in the entire record shows any lab, device sponsor or health-AI company buying clinical expert data from a vendor — not one contract, not one named supplier, not one rate.

high confidence8 minupdated 2026-08-30gaps · open questions · research agenda · diligence

Everything else in this dossier is an argument built on evidence of varying quality. This is where the evidence runs out, ordered by how much the answer would move the decision, each with the action that settles it. The general version is What we could not establish; nothing below duplicates it.


1. Has any lab ever bought clinical expert data from a vendor?

Nothing in the record says yes. OpenAI published a complete first-party recruitment funnel — 1,021 interest forms down to 262 compensated physicians — and named no supplier (HealthBench paper). Anthropic's healthcare launch names benchmarks and connectors, not a physician cohort. Every vendor in the vertical except Mercor claims lab relationships without naming one. The nearest thing to a counter-fact is OpenAI's own Human Data postings, which describe external vendors running "campaigns" — the same word its methods sections use (Ashby).

Why it is first. It is the difference between a market and an inference. Every revenue figure at The health read is built from hours and rate cards because no invoice exists anywhere in public.

What would settle it. Four conversations, none sensitive: Microsoft AI's Sr. Director, AI Data Acquisition and Operations, an open published role whose remit is vendor selection and negotiating "novel deals for data that have never been commercially available before" (microsoft.ai); OpenAI's Human Data programme managers; Vals AI, which assembled the licensor–benchmark–grader chain twice; and BioStack, which has a six-figure lab contract. Difficulty: low. Value: higher than anything else here.

2. Google DeepMind's AMIE clinician supply

AMIE is the most clinician-labour-intensive research programme at any lab and DeepMind has never described how it is staffed. OSCE-style evaluation requires simulated-patient actors, physician participants and physician graders — three separately paid populations — repeated across Nature (2025), Nature Medicine (2026), the disease-management paper (June 2026) and the video-consultation work (August 2026). The blog posts give no numbers, no recruitment method and no vendor.

Why it matters. If DeepMind sources patient actors and PCP raters through an agency, that agency is the closest existing analogue to the business and a live comparable. If it does it through institutional partnerships, that is a third procurement pattern alongside OpenAI's first-party funnel and Microsoft's Mayo route — and the one hardest for a vendor to compete with, because a Mayo partnership buys credibility no supplier can sell.

What would settle it. Read the full methods sections of the four papers, which were not reachable in this research. Failing that, MedGemma's model card names its dataset shapes — "US commercial biobank", "CRO", "European academic hospital", "US diagnostic center" — and those are identifiable with a few calls. Difficulty: low, and mostly reading.

3. Does Hippocratic AI pay its clinicians, and how much?

Its safety page states it "hired over 7.7K U.S. licensed clinicians to make 775K test calls" (Hippocratic), and its RWE-LLM framework describes 6,234 clinicians — 5,969 nurses and 265 physicians — evaluating 307,038 unique calls (Hippocratic). Neither document says whether they were paid.

Why it matters. This is the largest contracted clinician workforce documented anywhere in health AI, and it sits in Nursing, the niche this dossier ranks third. If they are paid, it establishes the market rate for nursing evaluation at genuine scale and the pricing model holds. If they are not — if clinicians will do this for access, credit and a byline — the pricing model does not survive contact with the largest buyer.

What would settle it. Ask three of the 5,969. The Nurse Advisory Council is public and its members are nameable. Difficulty: trivial.

4. Are Abridge's "expert clinical reviewers" staff, contractors or vendor-supplied?

Abridge plainly runs a reviewer pipeline — "structured feedback provided by expert clinical reviewers", annotations on de-identified encounters, physicians doing "blinded head-to-head comparisons" (tech.abridge.com) — and its own engineering blog explicitly declines to say what they are. It has 42 open roles and not one reviewer posting.

Why it matters. At a $5.3B valuation it is the largest ambient buyer, and the two readings point opposite ways. Vendor-supplied means a named incumbent exists in Ambient scribing that nobody has identified. Recruited off-platform through advisory boards and health-system partners means application companies are already doing the sourcing a vendor would sell them, informally and for free — which is the harder competitor.

What would settle it. Ask. The blog's authors are named. Difficulty: trivial.

5. What does an expert determination cost, and how long does it take?

No primary source gives either figure. The vendors who perform §164.514(b)(1) determinations — Datavant/Privacy Analytics, Mirador and peers — publish no price list [UNVERIFIED]. The working assumption of a low-five-figure engagement over one to three months, re-certified on a one-to-three-year cycle, is exactly that.

Why it matters. It prices the road not taken. Expert determination is a per-dataset, per-recipient, per-time-window cost with a specialist bottleneck — a structurally hostile shape for a business selling many small bespoke datasets to many buyers, and therefore an argument for the PHI-free architecture at Build it without ever touching a patient record. It is also quoted in every competitor's pitch, so you need the number even if you never pay it.

What would settle it. Two or three vendor quotes. Difficulty: low.

6. The EyePACS licence

Unverified, and it is the most common licence error in the field. EyePACS is the canonical public diabetic-retinopathy fundus corpus, it is in MedGemma's training mix, and it is best known through a Kaggle competition whose rules typically grant a licence for participation and academic research rather than unrestricted commercial redistribution. "It was on Kaggle" is not evidence of a permissive licence.

Why it matters. Dermatology and ophthalmology is ranked second in this dossier partly on corpus permissiveness. If EyePACS is restricted, the ophthalmology half of that argument leans entirely on de-novo material.

What would settle it. Obtain and read the actual terms before any use. Compare the honest cases: Stanford AIMI charges $70,000 per dataset per year for commercial access (AIMI), and 44% of the ISIC archive is CC-BY-NC and invisible unless you check per image. Difficulty: trivial, and non-negotiable before use.

7. Is this the practice of medicine?

No medical board opinion, statute or case squarely addressing whether writing cases or annotating data for AI training constitutes the practice of medicine could be found anywhere [UNVERIFIED]. The strong argument that it is not is that practice requires a physician–patient relationship and a specific patient, and this activity is closer to medical writing, examination-question authoring, expert-witness work or guideline drafting — all long-established things licensed clinicians do without it counting as practice.

Why it matters. It is the first question a careful recruit asks, and the answer decides whether cross-border contribution is viable at all. See What the doctor on the other end is risking.

What would settle it. A targeted legal memo in the two or three jurisdictions supplying most contributors — the cheapest piece of counsel in the whole plan. Difficulty: low, costs money not time.

8. Does any medical society have a position?

None found. Not the AMA, whose June 2026 policy is about AI supporting rather than replacing physician judgement [WEAK — located via news index, policy text not read]; not ACR, CAP, ASCO, AAD or AAO. On clinicians being paid to supply judgement as AI training data, organised medicine has said nothing.

Why it matters. A void is not a rejection, and it may be genuine white space. A business whose premise is paying physicians for judgement and putting it at the centre of AI development is aligned with organised medicine's stated concern rather than opposed to it. Being first to define the norm is cheaper than complying with someone else's later, and a specialty-society endorsement would solve recruitment and procurement in one move.

What would settle it. Approach one society with a methodology and an offer to co-publish. Difficulty: medium, and it is a business-development action rather than a research one.


The structural holes

Things nobody has written down, as distinct from things this research failed to find.

No lab has ever published a rate. Both HealthBench papers say only that "all members of the physician cohort were compensated". Mercor's published bands are the only first-party price points in the sector — and they are a cost input, not a sale price.

No per-artefact price exists for expert reasoning work. There are per-image, per-slide and per-study annotation prices, mostly from one vendor blog [WEAK], and per-hour expert rates. There is nothing on what anyone pays per rubric, per adjudicated disagreement or per reasoning trace — which is the unit you would actually sell. The only usable anchors are Mercor's ~$2,500 per benchmark task and Tempus's $81.25M for ~7M annotated slides (What actually gets sold).

No expert-data transaction in this space has a public price. Mercor/Sepal, Mercor/Deeptune, Datavant/Aetion, HealthVerity/Symphony — all undisclosed. The only priced deal in the neighbourhood is Datavant/Ciox at ~$7B, and that is a records business. There is no public comparable for what a clinical expert-data company is worth on exit.

The genetic counsellor wage is unretrievable, and it alone decides whether Genomics and variant curation is transformative or marginal. The NSGC survey is robots-disallowed and ACMG publishes no membership figure.

Pharmacovigilance has no LLM benchmark at all — a genuine unclaimed hole sitting in the one niche whose buyers do not buy data (Clinical trials and regulatory writing). And no commercial buyer for pharmacist-labelled data was found anywhere, only academic benchmark builders and two Mercor listings (Pharmacy).

No survey of clinician attitudes to selling judgement for AI training exists. Every claim in this dossier about what clinicians will accept is inferred from wage data and from the nursing union's opposition to AI replacing nurses. A twenty-clinician survey before launch would settle the pricing and consent-design questions at once.

The verdict those gaps sit under is The health read; the sequence for closing the top four is Ninety days in health.