Miju Labs

The security dossier

Build it without ever touching a patient record

MIMIC is closed to a commercial vendor on four independent grounds, and 44% of the ISIC archive is silently non-commercial — but the de-novo case route clears HIPAA, the Common Rule and institutional data ownership at once, and a JAMA trial in 2000 found written vignettes scored 71.0% against chart abstraction's 65.6% as a measure of real practice.

high confidence11 minupdated 2026-08-30phi · de-identification · mimic · licensing · isic · vignettes · architecture

This page and the three that follow are desk research by non-lawyers, compiled from primary legal sources on 30 August 2026. It is not legal advice. Before capital is committed, the documents listed at the end of What the doctor on the other end is risking must be read in full and run past counsel in each relevant jurisdiction. Not repeated after this.

The design question is usually asked backwards. Founders ask how to get lawful access to patient data, then build a compliance function around the answer. The better question is whether the product needs patient data at all — and for what frontier labs and device sponsors are actually buying, it does not.

The claim is specific: a corpus of de-novo clinical cases, adjudicated by a blinded panel, is not a privacy-constrained substitute for real records. It is a better instrument, and European law says so out loud. Four parts follow — the route that is closed, the route that works, the licence map for everything in between, and the architecture that falls out.

MIMIC is closed. Stop designing around it.

Nearly every health-AI data plan starts with MIMIC-IV or MIMIC-CXR, because they are free, large, real and famous. For a commercial vendor they are unusable — four independent blockers, each sufficient alone.

Both corpora are governed by the PhysioNet Credentialed Health Data License 1.5.0 (licence text):

In their words

"The LICENSEE will use the data for the sole purpose of lawful use in scientific research and no other."

"The LICENSEE will not share access to PhysioNet restricted data with anyone else."

Blocker one — the purpose limitation. "Scientific research and no other" excludes commercial product development and excludes producing labels for sale. Whether a given AI-lab engagement counts as scientific research is not a question to be arguing against MIT-LCP after the corpus is built.

Blocker two — the sharing prohibition kills the core operation. The business routes records to a panel of contracted clinicians. That is sharing access. Credentialing each annotator does not obviously fix it, and hands your supply pipeline to someone else's licensing queue.

Blocker three — the licence runs to a natural person. The parallel Data Use Agreement 1.5.0 is first-person singular throughout — "I will not share access," "I have requested access… for the sole purpose of lawful use in scientific research." There is no corporate licensee. An employee's credentialing does not licence the employer.

Blocker four — third-party AI APIs are explicitly out. PhysioNet's notice "Responsible use of MIMIC data with online services like GPT" (18 April 2023) confirms the DUA "explicitly prohibits sharing access to the data with third parties, including sending it through APIs provided by companies like OpenAI, or using it in online platforms like ChatGPT."

Resale of derived labels is not addressed explicitly, which is worse than a prohibition rather than better [UNVERIFIED] — silence in an agreement whose purpose limitation you are already outside. The commercial answer is the same either way: a label set keyed to MIMIC row identifiers is worthless to anyone not themselves credentialed.

MIMIC is a contamination risk, not just a closed door

The rule is not "do not build on MIMIC." It is do not let credentialed staff touch it on company time or company hardware. A corpus whose provenance ledger cannot rule out contamination from a research-only licence fails buyer diligence, and you cannot prove that negative after the fact.

The de-novo route, and why it is stronger than it sounds

The obvious objection to made-up cases is that they are made up. That objection was tested twenty-six years ago and lost.

Peabody et al., JAMA 2000;283(13):1715-22 (PubMed) ran three methods head to head across four common outpatient conditions at two VA medical centres, using unannounced standardised patients — trained actors presenting to real clinics — as the gold standard. Against that benchmark they compared chart abstraction of the same visits and physicians' responses to written vignettes corresponding exactly to the SP presentations.

The written hypothetical beat the medical record

Quality of care measured at 76.2% by standardised patients, 71.0% by vignettes, 65.6% by chart abstraction. Vignettes sat closer to the gold standard than the real chart did, and the pattern held when disaggregated by condition, by case complexity, by site and by level of physician training (P<.001 to <.05).

The mechanism is not mysterious. Chart abstraction measures documentation practice; the vignette measures reasoning. Documentation is a lossy, copy-forward-contaminated record of what a clinician actually thought. That reasoning is what an AI lab wants to train and evaluate against.

State the limits before a buyer finds them: quality-of-care process measurement, four common outpatient conditions, US VA primary care, 1997. It validates vignettes as a measure of physician practice, not as training data for models, and says nothing about imaging or rare disease. Do not over-claim it — it still points the wrong way for anyone arguing only real records are valid.

Three regimes fall away at once

No PHI. 45 CFR §164.514(a) states the standard positively: "Health information that does not identify an individual and with respect to which there is no reasonable basis to believe that the information can be used to identify an individual is not individually identifiable health information" (govinfo). A case about a person who does not exist relates to no individual. There is no de-identification step because there was never identification.

No human subject. 45 CFR §46.102(e) requires a living individual — a human subject is "a living individual about whom an investigator… obtains information" or "obtains, uses, studies, analyzes, or generates identifiable private information" (govinfo). A fictional patient is not one. No IRB.

No institutional data ownership. No hospital record was accessed, so no hospital has a claim on the derivative.

The real risks are contractual and epistemic, not statutory

Nothing above stops the business. What can stop it is a clinician writing up a remembered patient rather than a constructed composite — rarity is identifying, and "the 34-year-old ballet dancer with the aortic dissection last March" is PHI-equivalent content produced from memory. What can devalue it is realism drift: cases cleaner and more canonical than reality, over-representing what clinicians find teachable rather than what is common. Both are engineering problems with process answers — parameter grids the author does not choose, a rarity ceiling, a per-case attestation. See What the doctor on the other end is risking for the employment-contract half.

The affirmative case is stronger than the defensive one. De-novo cases are the only instrument permitting controlled variation: matched pairs identical but for one attribute, counterfactual probes, calibrated difficulty, coverage of presentations too rare to sample. You cannot randomise a real patient's demographics while holding the presentation fixed. Sell the capability, not the compliance — see What actually gets sold for that unit priced.

The corpus licence map, with the real numbers

Where real data helps, the constraint is licensing, and the numbers are worse than the field assumes.

CorpusLicenceCommercial trainingVerdict
MIMIC-IV / MIMIC-CXRPhysioNet Credentialed 1.5.0, per-personNo — "scientific research and no other"Closed
CheXpertNon-commercial by defaultOnly under paid licence, $70,000/dataset/yearExpensive, viable
ISICPer-image CC-0 / CC-BY / CC-BY-NCYes for ~307k of 553k imagesUsable, with filtering
ClinVarPublic, attribution requestedYesUsable
PMC Open Access SubsetPer-article CC; sanctioned bulk channels onlyYes for CC-BY/CC-0 itemsUsable, with filtering
EyePACS[UNVERIFIED]Unknown — assume restrictedVerify before touching

ISIC is the instructive case. Pulled live from the ISIC API on 30 August 2026: 553,019 images — 48,751 CC-0, 258,980 CC-BY, 245,288 CC-BY-NC. About 44% of the archive is commercially unusable, and the split is invisible unless you check per-image metadata. Anyone who bulk-downloaded "ISIC" and trained on it has probably breached CC-BY-NC without noticing.

CheXpert sets the price of the alternative. Stanford AIMI states its datasets are shared under a licence "restricted to non-commercial purposes," with commercial access available through application, committee review and an annual fee of $70,000 USD per dataset (FY25) (Stanford AIMI). An honest answer, and an anchor in both directions: a cost you would bear, and a comparable when pricing your own product.

ClinVar and PMC are genuinely open, and therefore carry no moat. NCBI asks only for attribution and notes that "NIH does not independently verify the submitted information" (ClinVar) — that admitted imperfection is a market for expert re-adjudication, which is the Genomics and variant curation play. PMC's OA subset is per-article, and carries a trap: "Systematic retrieval (or bulk retrieval) of articles through any other automated process is prohibited" outside the sanctioned channels (NLM).

EyePACS stays unverified until someone reads the terms. It is best known through a Kaggle competition, and "it was on Kaggle" is the single most common licence error in the field: competition rules typically licence participation and academic research, not commercial redistribution.

Licence-hygiene auditing is a product, not overhead

Every buyer with an image pipeline assembled before 2025 probably has CC-BY-NC contamination in it, invisibly, because the split lives in metadata nobody checked. A per-item licence provenance ledger is a defence for you and a saleable audit for them — the cheapest possible first conversation with a health-AI company about to face diligence. The law is about to arrive applies the same logic to how you engage the clinicians.

The two de-identification routes, and why you want neither

HIPAA offers two ways to make PHI not-PHI. Both are structurally hostile to this business.

Safe Harbor, 45 CFR §164.514(b)(2), removes eighteen categories of identifier. Item (R) — "any other unique identifying number, characteristic, or code" — is not a list but a standard, and it swallows most of what makes clinical data valuable: a rare disease, an unusual treatment sequence, a distinctive imaging finding. The date rule permits only the year, gutting any dataset whose value lies in the interval between events. And (b)(2)(ii)'s "actual knowledge" test is an ongoing obligation, not a one-time checklist — identifiers must go "regardless of location in a record," free text included (OCR guidance).

Expert Determination, §164.514(b)(1), requires "a person with appropriate knowledge of and experience with generally accepted statistical and scientific principles" to determine "that the risk is very small" and document the analysis. HHS is explicit that there is no credential requirement and "no explicit numerical level of identification risk" universally satisfies the standard. The determination is made for an anticipated recipient and an anticipated use, so it is not portable.

What a determination costs could not be established

Neither cost nor duration is verifiable from any primary source [UNVERIFIED] — the vendors who perform these determinations do not publish price lists. Get quotes before repeating any figure. The structural point survives the gap: expert determination is a per-dataset, per-recipient, per-time-window cost with a specialist bottleneck. For a business selling many small bespoke datasets to many buyers, that cost shape never amortises.

This bites hardest in Radiology and Genomics and variant curation. A renderable head CT is arguably a "comparable image" under §164.514(b)(2)(Q), so Safe Harbor forbids the data rather than cleaning it; a germline sequence carries no listed identifier and is nonetheless a permanent unique one, so item (R) arguably makes the sequence itself unremovable. In both, de-identification is not a gate you pass through — it is a permanent tax on utility.

The reference architecture

In build order.

1. Case origination — de-novo only. Clinicians author against a written specification. The parameter grid (age band, sex, comorbidity profile, setting, acuity, atypicality, guideline vintage) is generated by the company, not chosen by the author. Rarity ceiling with escalation review. Per-case attestation: composite, specification-driven, not a remembered individual, no institutional confidential information, no employer resources. The attestation is also the audit artefact you sell.

2. Public-corpus layer — filtered, audited, strictly optional. ISIC CC-0 and CC-BY subsets filtered per image via the API; ClinVar; PMC OA via sanctioned bulk channels. Hard institutional ban on MIMIC and any non-commercially-licensed corpus touching company systems or personnel. Per-item licence ledger from day one, because it cannot be reconstructed later.

3. Annotation — built to the FDA specification from the first item. Independent, mutually blinded assessment by at least three credentialed clinicians. Documented grading protocol, a record of exactly what each annotator saw, intra- and inter-rater variability retained, structured adjudication with recorded reasoning. Never collapse to consensus. The regulator wrote your product spec explains why this shape is worth more than the labels it produces.

4. Clinician engagement — the primary risk surface. Company-issued devices and accounts, personal time only, express IP assignment, warranty of no employment conflict, indemnity. What the doctor on the other end is risking has this in full.

5. Commercial layer — sell evidence packages, not rows. Data, reference-standard dossier, inter-rater analysis, adjudication log, credential register, licence ledger. Start with model-output review over synthetic cases: no PHI, no human subject, no covered entity, highest margin, buildable this quarter.

6. Deliberate exclusions, each a line in the sales deck. No business associate agreements. No Security Rule scope. No breach-notification exposure — above 500 individuals a breach becomes a permanent public entry naming your company, closer to extinction than to a fine for a trust business. No expert determinations, no re-certification treadmill.

The constraints that look like barriers are the reason the business exists. Europe prefers the design you were going to build shows Europe does not merely permit this architecture — Article 89(1) GDPR arguably requires it. Characterised disagreement and The health read are where the dossier lands.