Miju Labs

The security dossier

Ninety days in health

Ask the one question nobody has asked, seat twenty coders against a 1.2% funnel, and build a single 500-chart blinded panel — because the inter-rater statistic is the only deliverable that cannot be added afterwards. Every step below is ordered so that a no arrives early and cheap.

medium confidence8 minupdated 2026-08-30execution · recruiting · protocol · metrics · sequencing

Ninety days answers three questions and no others. Has anyone ever bought clinical expert data from a vendor? Can you seat twenty credentialed contributors against a 1.2% funnel? And will a buyer pay three times a single-label price for a blinded three-reader panel? Everything below is ordered so that a no arrives as early and as cheaply as possible.

The general version of this sequence is The first ninety days. The health-specific difference is that one deliverable — the inter-rater statistic — cannot be retrofitted, so the protocol has to be right before the first read rather than before the first sale.

Days 0–10: ask the question nobody has asked

The single highest-value action available is establishing whether any lab, any device sponsor or any health-AI company has ever bought clinical expert data from a vendor. Nothing in the record says yes. OpenAI published a first-party funnel and named no supplier; Hippocratic hired 7,700 clinicians itself; Abridge's engineering blog describes "expert clinical reviewers" and declines to say what they are (tech.abridge.com). See The labs as health buyers and Health-AI companies as buyers.

Four calls settle it, and none of them is a sensitive request. Microsoft AI's Sr. Director, AI Data Acquisition and Operations is a published, open role whose remit is to "identify critical data needs, select vendors, and personally negotiate commercial agreements" and to structure "novel deals for data that have never been commercially available before" (microsoft.ai) — that seat is identifiable and it is literally the buyer. OpenAI's Human Data function has three open roles and a stated interface to "external vendors" (Ashby). Vals AI has assembled the three-party chain of licensor, benchmark house and credentialed graders twice and knows what each leg cost. And BioStack has a six-figure lab contract and seven employees (BioStack and Sepal).

Ask two things in the same conversations. What did you pay per adjudicated item? And would you pay three times a single-label price for three blinded readers plus an adjudication log?

What it disproves. If the answer to the second question is no across four buyers, the argument at Characterised disagreement is dead and the remaining business is single-annotator labelling, which is a Hyderabad business rather than a European one. That is a ten-day finding and it saves a company.

Days 0–15: set the architecture before anyone is hired

Two decisions are irreversible in practice and both are cheap now.

PHI-free is a product architecture, not a compliance posture. Never acquire, store or transmit data about a real identifiable patient. The route is de-novo authoring against a company-generated parameter grid — age band, sex, comorbidity profile, setting, acuity, atypicality — with a rarity ceiling, escalation review, and a per-item attestation that the case is a specification-driven composite rather than a remembered individual. That attestation is simultaneously the control and an audit artefact you sell. The evidence base is Peabody's 2000 JAMA trial, in which written vignettes scored 71.0% against chart abstraction's 65.6% measured against unannounced standardised patients as gold standard (PubMed) — the written hypothetical beat the medical record. Full argument at Build it without ever touching a patient record.

Ban MIMIC on company systems and personnel from day one, including on personal research time. The PhysioNet Credentialed licence runs to a natural person, permits use "for the sole purpose of lawful use in scientific research and no other", and forbids sharing access with anyone else (PhysioNet). The risk is not the breach; it is corpus contamination, which is unprovable afterwards and fatal in diligence. Where public images are genuinely useful, filter ISIC per image via its API — of 553,019 images, 48,751 are CC-0 and 258,980 CC-BY, while 245,288 are CC-BY-NC, and the split is invisible unless you check. Maintain a per-item licence provenance ledger.

Days 5–25: seat twenty, against the real funnel

Recruit in Medical coding, not in a physician specialty. Twenty AHIMA CCS or AAPC CPC holders at the observed $80–110/hr band, verified for free through both bodies' credential lookups, sourced through AAPC local chapters, AHIMA component state associations, the AHIMA Career Assist job bank and r/MedicalCoding.

Plan against 1.2%. That is what emailing 12,000 physicians from IQVIA's OneKey database converted — 149 completers (arXiv). Coders should convert better because the wage multiple is 3x rather than 1x and displacement anxiety works in your favour: autonomous coding is visibly eating entry-level work, so a credentialed coder is motivated to take AI-adjacent work. But do not assume it; measure it. The registry detail is at What a clinician hour costs.

Two contractual items go into onboarding, not into a later policy document. A warranty that the work does not breach the contributor's employment terms, with a hard rule of personal time, company-issued accounts and no employer materials — employer IP is the largest practical risk in this business and it is the one a Series A diligence finds (What the doctor on the other end is risking). And company-carried professional and technology E&O with contributors indemnified, because standard malpractice cover almost certainly does not extend to work with no patient.

What it disproves. If twenty credentialed coders cannot be seated in three weeks at $80/hr, the reach: 5 at The health read is wrong and every downstream cost estimate is optimistic.

Days 15–45: build one blinded panel, and fix the protocol first

Produce one 500-chart adjudicated evaluation set in ICD-10-CM coding. De-novo charts to the parameter grid, then the part that matters:

  1. Three credentialed coders read each chart independently and blinded to one another. Record exactly what data each reader saw.
  2. Compute and retain inter-rater agreement per task before any adjudication happens.
  3. Adjudicate disagreements by a fourth reader, with the reasoning recorded — not just the resolved code.
  4. Never collapse to consensus. Store every independent assessment, with annotator identity and credentials, permanently.
  5. Write the equivocal-case policy and the incorrect-annotation remediation plan before the first read, not after the first complaint.

That list is a direct implementation of FDA's January 2025 draft guidance, which asks for annotator expertise, blinding, adjudication method and "an assessment of the intra- and/or inter-clinician variability for each task" as marketing-submission content (FDA). Step 2 is the one that cannot be added later: if readers were not blinded when they read, the statistic does not exist and cannot be manufactured. See The regulator wrote your product spec.

Budget roughly 0.8–0.9 expert-hours per chart across three reads and adjudication, so $50,000–$70,000 of expert labour for the set [INFERENCE]. Hold back a second, identically constructed set and never release it.

What it disproves. If three certified coders agree on 95% of charts, there is no disagreement structure to sell and the panel product collapses into a single-label product. The pathology precedent says otherwise — eight dermatopathologists agreed unanimously on only 53.5% of 792 melanocytic cases (Nature Communications) — and Vals's own MedCode used two coders per sample and published no κ (Vals). But coding is more rule-bound than melanocytic pathology and the agreement could be high. This is the experiment, and it is the only one that matters.

Days 45–75: publish the number, then convert

Publish two things together: the inter-rater agreement statistic for certified coders on ICD-10-CM assignment — which nobody has published — and frontier model scores against the held-out set, positioned against MedCode's public 63.57% ceiling.

Then work three buyer types in parallel, because they buy differently. Health-AI application companies are the likely first cheque: CodaMetrix, SmarterDx, Fathom, Arintra, Nym, AKASA, and Ambience, which has already run an 18-physician, four-clinician-consensus gold panel and therefore already believes in the product (Health-AI companies as buyers). Device sponsors are the higher-price, slower buyer, and the pitch is the dossier rather than the data — target the Modification Protocol inside a Predetermined Change Control Plan, which commits the sponsor to re-running validation against your specification for the life of the device. A lab is the reference logo, and no health company has ever obtained one.

Day 90: the numbers that mean continue or stop

ContinueStop
Day 30Four buyer conversations held; at least one says it would pay a premium for blinded multi-reader work. Twenty credentialed coders seated and verified. Architecture decisions written down and enforced technically.No buyer will price a panel above a single label. Fewer than ten contributors seated at the advertised band.
Day 60500 charts triple-read and adjudicated for under $70k. Unanimous agreement below ~85%, i.e. a real disagreement structure. Model scores reproduced against MedCode's public band.Agreement above 95%, so there is nothing to characterise. Cost per adjudicated chart above $200, which kills the margin.
Day 90One paid pilot signed at or above $50,000, or two at $25,000. A published κ number that got inbound. A named second-round conversation with a device sponsor about a PCCP.Ninety days of interest and no purchase order. Every buyer asking for cheaper single labels — which is the trap at Characterised disagreement, and taking it is worse than stopping.

One deliberate omission. Nothing above builds an imaging pipeline, signs a hospital data agreement or applies for a controlled-access corpus. Each of those turns a clinician-recruitment company into a hospital-business-development company, and there is no ninety-day version of that. The open items this sequence does not close are at What the health dossier could not establish.