Miju Labs

The security dossier

Medical coding

The best labour arbitrage in the dossier and probably the best niche of all thirteen: a 3x wage spread, the only pool with first-party date-stamped credential counts, no licensure, no malpractice, no PHI, and the best model on earth scoring 63.57%.

buildhigh confidence8 minupdated 2026-08-30

A certified professional coder costs about $29 an hour at market. OpenTrain AI is advertising a Medical Coding AI Evaluation Lead at $80/hour, remote, US-only, 20+ hours a week (OpenTrain). Mercor lists a risk-adjustment/HCC coding leader at $110/hour and an HIM coding leader at $80/hour (Mercor). That is a 2.6x to 3.6x spread on the cheapest credentialed labour pool anywhere in healthcare.

Now the capability gap. On Vals AI's MedCode leaderboard — ICD-10-CM assignment from de-identified discharge summaries, 2,755 codes, each sample independently annotated by two certified coders — the best model in the world is Claude Opus 5 at 63.57%, ahead of Gemini 3.1 Pro Preview at 59.06% and Claude Fable 5 at 56.07%, across 86 models (Vals AI, 20 August 2026). Eighty-six frontier models are publicly bad at this, and everyone can see the scoreboard.

And there is no licence to hold, no malpractice to carry, no scope of practice to breach and no protected health information to touch. Nothing else in this dossier has all four.

Correct the 88-vs-56 figure

The circulated "88% note generation versus 56% diagnosis coding" pair is a real comparison — Vals MedScribe against Vals MedCode — but an older snapshot. As of 20 August 2026 it is ~91% against ~64%, and 56.07% is now the third-placed model on MedCode. The durable fact is the ~27-point gap, not the levels.

What the artefact is

Four things, and the fourth is the one nobody supplies.

  1. Gold-standard code assignments — ICD-10-CM/PCS, CPT/HCPCS and DRG on de-novo or de-identified charts, dual-coded by two credentialed coders with documented adjudication of disagreements.
  2. HCC and risk-adjustment capture with MEAT-criteria justification for each condition claimed.
  3. Denial-to-appeal pairs — the denial, the appeal letter that overturned it, and the payer policy citation that did the work.
  4. Rationale traces. The guideline text the coder applied, why it applies, and why the near-miss code was rejected. This is the highest-value and least-supplied artefact here, because it turns a label into a training signal, and because a coder produces it naturally: rejecting the plausible-but-wrong code is the job.

The structural point is that coding correctness is rule-defined — fixed by the ICD-10-CM Official Guidelines, AHA Coding Clinic, CPT Assistant and the NCCI edits, not by any particular patient. That is what makes the fourth artefact writable and the whole niche cheap.

Does it need patient data

No, and this is the niche's structural advantage rather than a workaround.

A CPC can write a de-novo operative note, progress note or discharge summary and its correct code set from scratch. Writing a plausible chart is a core coder skill — it is what they read all day. Because correctness is defined by a published rule set rather than by a real encounter, the fictional chart is a complete substrate, not a degraded one. There is no de-identification step because there was never identification, no human subject under the Common Rule, and no institutional data-ownership claim because no institutional record was opened. See Build it without ever touching a patient record.

Training coding models on privacy-preserving synthetic clinical data is already a published method (arXiv 2603.23515), so a buyer's diligence question here has a citation rather than an argument.

The real exposure is employer IP: a coder must not bring charts from their day job. De-novo authoring is the mitigation and it is the natural mode for this pool. Enforce it with company-issued accounts and a per-chart attestation, and the attestation becomes an audit artefact you sell alongside the labels.

Is anyone buying

Budget scores 4 — not the 5 that Clinical reasoning earns, because the disclosed cheques come from health-AI application companies rather than the labs, and no first-party lab posting anywhere names a CPC. But the buyer bench is crowded:

BuyerMoneyEvidence
Ambience Healthcare$243M Series C, $1.25B valuationICD-10 model benchmarked against 18 board-certified physicians; each encounter labelled by a consensus of four or more expert clinicians; 27% relative improvement over the physician baseline (Ambience, CNBC)
CodaMetrix$40M Series B, later $55M(MobiHealthNews)
SmarterDx$50M(SmarterDx)
FathomCVS Health Ventures strategic investment, May 2026(HIT Consultant)
Arintra$25M, Aug 2026[WEAK — trade coverage via news index]

Plus Nym Health, AKASA and Martlet AI, plus the provider side: autonomous coding lifted revenue 5.1% at Mercyhealth (Healthcare IT News, 10 April 2026).

The frontier labs are an indirect buyer, but a real one. MedCode publicly scores 86 models on ICD-10 assignment and they are all losing. A benchmark that every lab is visibly bad at is the best sales artefact an expert-data company can hold, and it costs nothing to point at. See Health-AI companies as buyers and The measurement landscape.

Payers are the third leg and the slowest: CMS's intensification of RADV audits drives the risk-adjustment side, but it is a procurement-heavy sales cycle — large and recurring, not the first cheque. See Pharma, payers and providers.

Proof scores 4. You can be visibly the best source here faster than anywhere else in the dossier: a public leaderboard to beat, an adjudication protocol you can publish, and no incumbent vendor to displace.

What the expert costs

Cost scores 5, the joint-cheapest pilot in the dossier alongside Clinical reasoning — and cheaper in absolute labour terms by a factor of five.

  • BLS, Medical Records Specialists (SOC 29-2072), May 2025: median $51,140/yr = $24.59/hr; 200,700 jobs; +8% projected 2025–2035, "much faster than average" (BLS).
  • AAPC's 2025 salary survey (2024 data): CCA $48,321; CCS $57,500; CPC $59,605; two credentials $71,130; three or more $76,035 (secondary transcription) [WEAK — AAPC's own survey page 403s to automated fetch].

That is $22–30/hr at market against $80–110/hr observed. Compare the physician niches, where the observed AI rate sits below clinical opportunity cost — see What a clinician hour costs for why the arbitrage runs backwards in radiology and forwards here.

No equipment, no data-use agreement, no institutional relationship, no scanner. The first pilot is twenty credentialed coders, a rule-set spec, a dual-coding harness and an adjudicator.

Getting to them

Reach scores 5, and it is the best-evidenced 5 in the dossier because the pool is counted.

AHIMA publishes active credential-holder totals, per credential, date-stamped to 31 December 2025: CCS 36,925 (2025 pass rate 84% on 6,331 first-time testers, AHIMA); CCA 7,753 (62% on 2,101, AHIMA); RHIT 26,128 (AHIMA); CDIP 2,913 (AHIMA). No other pool in this dossier has first-party per-credential counts with a date on them — nursing's national licence count is unretrievable, pharmacy's does not exist, Genomics and variant curation has no membership figure at all.

AAPC reports "over 300,000 worldwide members" as of September 2025 (via Wikipedia) [WEAK], across CPC, COC, CIC, CRC, CPB, CPMA and 20+ specialty credentials (AAPC).

Both bodies run credential-verification lookups, so quality gating costs effectively nothing per candidate. Named channels: AAPC local chapters, AHIMA component state associations, the AHIMA Career Assist job bank, r/MedicalCoding, the AAPC member forums, and RISE National for risk adjustment.

One recruiting note that is unusual and useful: displacement anxiety works in your favour here. Autonomous coding is visibly eating entry-level work, so a credentialed coder is motivated to take AI-adjacent work at two to three times their day rate. Contrast Nursing, where the union's position makes the same pitch dangerous.

Where the benchmarks sit

BenchmarkStatusFrontier score
Vals MedCode (ICD-10-CM, 2,755 codes, 86 models)Wide openOpus 5 63.57%; Gemini 3.1 Pro Preview 59.06%; Fable 5 56.07%
Vals MedScribe (SOAP notes, 87 models)Near-saturated — the contrast caseOpus 5 90.99%
MedHELM — medical codingOne dataset for the whole categoryMIMIC-IV Billing Code, gated
MedHELM — Administration & WorkflowWeakest category in the suite0.53–0.63
MAX-EVAL-11 (ICD-11, Oct 2025)New, unclaimed
NEJM AI, "LLMs Are Poor Medical Coders"Canonical negative result[WEAK — full text 403s]

Sources: Vals MedCode, Vals MedScribe, MedHELM, MAX-EVAL-11, NEJM AI.

Two readings. MedHELM's 121-task taxonomy contains exactly one medical-coding dataset, and it is gated — a category-shaped hole in the field's structural map. And models "struggle significantly with mental-health diagnoses" on MedCode: a named, addressable failure mode that a targeted 500-chart set could attack directly.

What would kill it

Defense scores 4 and room scores 4.

The pool is not scarce. Three hundred thousand AAPC members and no licensure barrier means nobody corners this supply. What holds the position is churn: ICD-10-CM updates every October, CPT every January, Coding Clinic and the NCCI edits quarterly. A coding eval set decays on a published schedule. That is a subscription, and it is why defense is 4 rather than 2.

The real kill is a buyer building it. Ambience has already run an 18-physician, four-clinician-consensus gold panel; CodaMetrix, SmarterDx and Fathom all employ coders. The counter is that such panels are episodic and project-shaped — Ambience's was built for one model release — which is precisely the shape a specialist absorbs better than an internal team.

The second kill is model progress. If MedCode goes from 63.57% to 92% in eighteen months the way MedScribe did, the commercial logic shifts from producing ground truth to producing harder cases and disagreement signal: a smaller, differently priced product. Nobody has published evidence of the shift yet, but MedScribe is the worked example of how fast it can happen — see Ambient scribing.

No first-party lab posting names a CPC

Lab-side demand here is inferred from Mercor's rate card and one OpenTrain listing. No posting from OpenAI, Anthropic, Google DeepMind or Microsoft AI names a certified coder — or any allied-health credential at all, the same hole that applies to RNs and PharmDs across this dossier. Mercor does not list roles it is not filling and OpenTrain's $80/hr is a real advertised band, so the demand is not imaginary. But it is intermediary-observed, not lab-disclosed. See Expert data for frontier labs.

Where the record is thin

AAPC's own numbers are the soft spot in an otherwise hard evidence base. The 300,000+ membership figure is a third-party citation of a September 2025 source, and AAPC's salary-survey page 403s, making the CPC $59,605 figure a secondary transcription. AHIMA's counts are first-party and precise; AAPC's are not.

Nothing is known about unit pricing. We have hourly rates from two sources and volumes from none. What a buyer pays per adjudicated chart, per rationale trace or per resolved dual-coding disagreement is unpublished everywhere, so any revenue model here is built from hours, not invoices.

The payer leg is asserted rather than observed: RADV pressure is documented, but no payer has been shown purchasing coder-labelled evaluation data from anyone. Compare The health read and Characterised disagreement.