Miju Labs

The security dossier

Clinical trials and regulatory writing

Where the dossier's own buyer thesis breaks: every disclosed pharma-AI deal from 2020 to 2026 is compute, cloud or platform, none is expert labelling — and freelance regulatory writers already earn $131.65/hr, so there is no arbitrage to capture.

avoidhigh confidence8 minupdated 2026-08-30

Every disclosed big-pharma AI deal from 2020 to 2026 is compute, cloud, platform access or drug discovery. None is expert labelling, data annotation or evaluation-set procurement.

PharmaPartnerValueDateWhat was bought
MerckGoogle Cloud$1B multiyearApr 2026Gemini Enterprise deployment
Eli LillyNVIDIA$1B over 5 yearsJan 2026Co-Innovation Lab, BioNeMo, GPUs
Eli LillyInsilico Medicine$2.75B milestonesMar 2026AI small-molecule design
Novo NordiskOpenAIundisclosedApr 2026GenAI across R&D, manufacturing, supply chain
Bristol Myers SquibbAnthropicundisclosedMay 2026Enterprise-wide Claude adoption
BayerGoogle CloudundisclosedApr 2024Radiology AI
ServierGoogle CloudundisclosedJan 2025Value-chain AI

The tracker that compiled these says it in as many words: "None of the disclosed deals explicitly mention expert-labeling, data annotation, or evaluation sets" (IntuitionLabs). Even the one deal that sounds like an exception is not one — GSK backing Relation Therapeutics with $110M for a "biological data factory" is wet-lab data generation, not expert judgement.

Then the second problem. Freelance regulatory writers focused on pharmaceutical work average $268,847 gross and $131.65 an hour (Clinicians Guide, reporting AMWA's 2024 survey) [WEAK — AMWA's own survey data is member-gated]. Employed regulatory writers average $166,457; biotech $198,322; pharma $192,619.

A niche whose experts already earn $131.65/hr in their day job has no labour arbitrage. The AI-data premium is at best zero and plausibly negative. This is the worst cost structure of the thirteen, which is why cost scores 1 — the only 1 outside Radiology.

What the artefact is

Six things, well-specified and genuinely valuable — which is the frustrating part:

  1. Protocol-design critique with ICH E6(R3)/E8 and FDA-guidance citations.
  2. CSR section drafting and gold-reference comparison.
  3. Safety-narrative writing from case data.
  4. Pharmacovigilance ICSR case-processing labels — seriousness, expectedness, causality against WHO-UMC or Naranjo.
  5. Eligibility-criteria structuring and patient-matching adjudication.
  6. Regulatory-submission Q&A against published guidance.

Item 4 is the one with a real hole behind it. There is no LLM benchmark anywhere for ICSR causality, seriousness or expectedness assessment, and the commercial PV-automation market is active (Datafoundry). It is a genuine unclaimed benchmark opportunity sitting inside the one niche whose buyers do not buy data.

Does it need patient data

Mostly no.

Protocol, CSR and regulatory work: no. These are built from public FDA and EMA guidance, ClinicalTrials.gov registrations, published CSRs and de-novo writing. There is no patient, no chart and no institutional record. The position is as clean as Medical coding's and for a similar reason — correctness is defined by published rule sets rather than by any individual.

Pharmacovigilance case processing: yes-ish. It needs realistic case narratives, which in principle means real adverse-event reports. In practice synthetic ICSRs are already standard industry practice, so the artefact is buildable without PHI, with the usual caveat that synthetic cases are cleaner than real ones and a buyer will say so.

Employer IP is severe here, and it is the operative constraint rather than PHI. Regulatory writers work on confidential submissions under strict NDAs, often naming the compound and the indication. Anything a contributor produces must be unambiguously de novo, and the risk is not that they copy a document — it is that a protocol critique for a fictional oncology trial reads uncomfortably like the one they wrote last quarter. Enforceable with discipline: company-issued accounts, personal time, per-item attestation, and a policy of declining contributors whose employment contracts cannot be cleared. See What the doctor on the other end is risking and Build it without ever touching a patient record.

Is anyone buying

Budget scores 2, and the composition of that 2 is the point of this page.

Pharma: no. Beyond the deal table above, Anthropic's healthcare announcement names Medidata as a clinical-trial data provider — that is trial operational data, not expert judgement. No payer has been found purchasing expert-labelled evaluation data either. See Pharma, payers and providers.

The labs: yes, and this is where the money actually is. Mercor's published bands price exactly this expertise, well above the pharma-services labour rate:

  • Disease-Area Clinician — Trial Endpoints & Prescribing: $150–230/hr
  • Medical Safety Expert: $140–190/hr
  • Medical Writer: $90–150/hr
  • Clinical / biomedical / pharma Evaluator: $80–120/hr

(Mercor)

That is a lab-funded market for regulatory and trial expertise, and the buyer is OpenAI or Anthropic or Google building pharma-vertical capability — not Pfizer's procurement department. Which reclassifies the whole niche. If you sell here, you are running a lab sale with a pharma-shaped product description, and you should price, position and forecast it as a lab sale. See The labs as health buyers and Expert data for frontier labs.

Proof scores 3. There is no commercial leaderboard to top, so becoming visibly the best source means publishing the PV benchmark nobody has built — achievable, credible, and watched by an audience that is not currently spending.

What the expert costs

Cost scores 1. This is the structural finding and it does not improve on inspection.

  • BLS, Technical Writers, May 2025: median $90,390/yr = $43.46/hr; 46,400 jobs; +1% 2025–2035 (BLS). This is the wrong series — it does not describe regulatory writers — but it is the only BLS series that covers them at all.
  • AMWA 2024 survey, secondarily reported: employed regulatory writers $166,457; biotech $198,322; pharma $192,619; full-time pharmaceutical-regulatory freelancers $268,847 gross, $131.65/hr (Clinicians Guide) [WEAK]. AMWA confirms a survey exists and is member-only; the publicly described version is the 2019 survey with 1,400+ respondents (AMWA).

Against Mercor's Medical Writer band of $90–150/hr, the arbitrage is negative at the midpoint. The only bands that clear the freelance rate — Medical Safety Expert at $140–190 and Disease-Area Clinician at $150–230 — are physician bands, not writer bands, which means the profitable version of this niche is buying physicians and the unprofitable version is buying writers.

Add the sales cost. Pharma procurement runs 9–18 months, requires vendor qualification, GxP and CSV documentation, and typically a master services agreement. It is the slowest sales cycle in this dossier and the least likely to reward a small specialist. The lab channel is faster, but the lab channel is Clinical reasoning's channel with a different label on it.

Getting to them

Reach scores 4. The bodies exist and run directories; the pool is small.

AMWA (American Medical Writers Association), RAPS (Regulatory Affairs Professionals Society, which awards the RAC credential), DIA, ACRP and SOCRA for clinical research coordinators (CCRP and CCRC credentials), and TOPRA in Europe. All run member directories and annual conferences, and RAC, CCRP and CCRC are verifiable credentials.

What is missing is a register. There is no NABP Verify, no Nursys, no AHIMA credential lookup covering the whole pool — verification is credential-body by credential-body, and the pool is an order of magnitude smaller than Medical coding's or Nursing's. Reach holds at 4 because the directories are good and the conferences are concentrated, not because the supply is large.

One genuine advantage: no licensure and no malpractice exposure. A regulatory writer is not a licensed professional in the state-board sense, so there is no scope-of-practice question, no state-by-state patchwork and no indemnity structure to build. It is the one thing this niche has entirely in its favour.

Where the benchmarks sit

BenchmarkStatus
TrialGPT and its real-world eligibility-screening adaptationPublished, academic (JAMIA 33(4):909)
LLMs for Clinical Trial Protocol AssessmentsClinical Pharmacology & Therapeutics, 2026 (DOI) [WEAK on contents — 403s]
Multimodal trial patient-matching pipelineReal-world validated (PMC)
MedHELM — Medical Research Assistance0.65–0.75; covers literature research, data analysis, quality assurance, enrolment
PharmacovigilanceNo LLM benchmark exists

Thin, academic, and with no commercial leaderboard — nothing here plays the role Vals MedCode plays for Medical coding. Room scores 4 on the strength of the PV hole, with the caveat that room is only worth having when someone is paying for it. See The measurement landscape.

What would kill it

It is already dead in its pharma form. What would kill the surviving lab-facing form:

Buyer behaviour is the primary blocker, not competition. Nobody is defending this position because nobody is holding it. The negative finding above is not "pharma buys from someone else" — it is "pharma buys a different category of thing."

The wage floor is the second. No vendor margin survives between a $131.65/hr expert and a $90–150/hr sell price. The only escape is to sell the physician-band artefacts — safety expertise, trial endpoints, prescribing judgement — which puts you back in Clinical reasoning's market competing with Mercor, Handshake and Surge.

Defense scores 4 on the two things that do hold: the RAC/CCRP/CCRC pool is small and credentialled, and regulatory knowledge refreshes with every new guidance document, ICH revision and precedent submission. If a buyer appeared, the position would be genuinely hard to copy. That conditional is doing all the work in the sentence.

The incumbent vendors who would in principle need evaluation sets — Certara CoAuthor (Certara) and Yseop (Yseop) — show no evidence of buying external expert data [UNVERIFIED either way].

The one large priced market here is not an AI market

Pharmacovigilance outsourcing was $7.24B in 2025 rising to $8.51B in 2026 and forecast to $16.09B by 2030 (The Business Research Company); medical affairs outsourcing adds roughly $3B. Together that is $11–12B of pharma paying for structured clinical review — case-level adjudication with audit trails — because regulators require it. It proves willingness to pay and it demonstrates the operating model. It does not demonstrate that any of that budget will move to AI evaluation data, and no bridge between the two has been observed. Whether the PV outsourcers themselves (IQVIA, ICON, Accenture, Qinecsa) become buyers of eval sets as they automate is the open question this niche turns on, and it is unanswered.

Where the record is thin

Whether Certara, Yseop or any regulatory-writing vendor buys external evaluation sets is unestablished in both directions — not found, not excluded.

AMWA's compensation data is member-gated, so the $131.65/hr figure that carries this page's central argument is a secondary transcription of a survey nobody outside AMWA can read. It is directionally certain — regulatory writing is well paid — but the precise number should not be quoted in a pitch without the primary source.

And the lab-channel reclassification rests on Mercor's rate card alone. Four priced bands prove labs are buying trial and regulatory expertise through an intermediary; they do not establish volume, cadence or whether the projects recur. Compare the ranked read in The health read and the entry argument in Characterised disagreement.