Two published studies show ambient-scribe companies buying external clinician review, which is more direct evidence of purchasing than any other niche in this dossier produces.
Suki published a validated evaluation of its own ambient scribe in Frontiers in Artificial Intelligence: 97 de-identified audio recordings across five specialties, 194 notes, 388 paired reviews, with two board-certified clinicians recruited per note, blinded to origin as "Model 1" and "Model 2." On PDQI-9 the gold note scored 4.25 against the ambient note's 4.20 (p=0.04); gold won on accuracy and succinctness, ambient on thoroughness and organisation; and hallucinations appeared in 20% of gold notes versus 31% of ambient notes (p=0.01) (Frontiers). Note the word recruited. Suki did not use its own staff. That is a purchase.
GraiLabs did the same in neurosurgery: 49 encounters, two consultant neurosurgeons scoring independently plus a third adjudicating the ten highest-disagreement cases. Omissions ran 15 for the AI against 131 for handwritten notes; notes free of clinically meaningful errors were 25/49 for AI against 4/49 for handwritten (Research Square).
So why is this the page with an avoid stance? Because of what happened to the benchmark.
The circulated "88% note generation versus 56% diagnosis coding" pair is Vals MedScribe against Vals MedCode, and it is an older snapshot. As of 20 August 2026 the numbers are ~91% (Claude Opus 5, MedScribe, 90.99%) against ~64% (Claude Opus 5, MedCode, 63.57%), and 56.07% is now the third-placed model on MedCode (MedScribe, MedCode). The ~27-point gap is the durable fact, not the levels — and the gap has held while the note-generation side climbed three points into saturation. That is the whole argument for preferring Medical coding to this page.
What the artefact is
Four things, in descending order of value and ascending order of difficulty:
- Blinded pairwise ratings of AI note against reference note on a validated instrument — PDQI-9 is the one the literature uses.
- Hallucination, omission and distortion adjudication with severity grading. GraiLabs' taxonomy — hallucinations, distortions, omissions, major clinical impact — is the working shape.
- Note-to-encounter fidelity annotation: every clinical claim in a note traced to a specific transcript span. This is the highest-value artefact and the one that needs real encounter data.
- Rubric authorship — what a specialty-correct SOAP note must contain. MedScribe's rubrics were authored by medical scribes and physician assistants who independently annotated each transcript, with conflicts reconciled collaboratively (Protege/DataLab). This is the one task in the niche that is not physician-priced.
Does it need patient data
Partially, and this is the only niche in the seven where the honest answer is not a clean no.
The state of the art already runs on synthetic transcripts. MedScribe generates them via an adapted NoteChat framework from real, de-identified SOAP notes, and validates the realism by A/B test: medical experts identified the true transcript only 39% of the time, against a chance rate of 50% and p > 0.05 (Vals). Experts could not tell. That is a strong result and it means the evaluation layer is buildable without touching PHI.
But the artefact that carries the most value — real encounter to real note fidelity — needs either a covered-entity BAA or a de-identified corpus under licence. Real encounter audio is PHI, and voice is a Safe Harbor identifier in its own right. So the PHI-free architecture in Build it without ever touching a patient record gets you the rubric and the rating layer, and stops short of the thing buyers most want validated.
Employer IP bites harder here than elsewhere too: a scribe or CDI specialist reviewing notes they wrote at work is a direct conflict, not a theoretical one. De-novo and synthetic material sidesteps it, at the cost of the representativeness argument.
Is anyone buying
Budget scores 4. The money is real, well-capitalised and documented — it is just being spent in a shape that suits an internal team.
| Company | Money | Buy or build |
|---|---|---|
| Abridge | $300M Series E at $5.3B, plus a $316M extension | Builds, with an unspecified reviewer pool (Fierce) |
| Ambience | $243M Series C at $1.25B; OpenAI Startup Fund, a16z, Kleiner | Clinical strategy in-house — two salaried clinical roles, no external reviewer network advertised (careers) — but buys gold panels project-by-project |
| Suki | Series D, $168M total | Buys — documented. Two board-certified clinicians recruited per note |
| GraiLabs | — | Buys — documented. Three consultant neurosurgeons |
| Microsoft/Nuance | — | Builds and publishes: ACI-Bench is a Microsoft/Nuance artefact, and it is public (arXiv 2306.02022) |
Ambience is the instructive case. Its careers page shows a build posture, yet its ICD-10 model was "benchmarked against 18 experienced, board-certified physicians," with each encounter "meticulously labeled by a consensus of four or more expert clinicians," producing a 27% relative improvement over the physician baseline (Ambience). So the honest read is: strategy in-house, gold panels bought episodically. That is a sellable engagement — high-spec, small-N, project-shaped — but it is not a subscription. See Health-AI companies as buyers.
Proof scores 2. With ACI-Bench public and MedScribe published, there is no leaderboard left to top and no obvious way to become visibly the best source in a domain whose leaders build their own.
What the expert costs
Cost scores 2 — the second-worst in the dossier after Radiology, and for the same reason.
The published studies used board-certified physicians, fellows, residents and advanced practice providers, two per note in Suki's case and three consultant neurosurgeons in GraiLabs'. Mercor prices RN and scribe reviewers at $55–65/hr but physician raters at $130–250/hr (Mercor). The rating work that buyers actually commissioned is at the physician end of that range.
That is the economic weakness of the whole niche, and it is worth stating as arithmetic. A blinded pairwise design needs two raters per note plus an adjudicator on disagreements — so roughly 2.2 physician-hours of cost per hour of rated output, against a sell price constrained by what an ambient-scribe company will pay for an episodic validation study. Compare Medical coding, where the same adjudication structure runs on $22–30/hr labour.
The rubric-authorship layer is the exception. MedScribe's rubrics came from scribes and PAs, not physicians, and that is genuinely cheap work. It is also a one-time build per specialty, not a recurring volume.
Getting to them
Reach scores 4.
The physician channels are the same as Clinical reasoning's and just as good — but the useful pool is specialty-matched, and that is the constraint. Suki needed five specialties; GraiLabs needed consultant neurosurgeons specifically. A general physician roster does not serve this niche; a roster indexed by specialty and by documentation-heavy practice setting does.
The cheaper half of the supply is easier and under-used: medical scribes, physician assistants and CDI specialists, reachable through the same HIM channels that serve Medical coding — AHIMA component associations and the Career Assist job bank. That population authored MedScribe's rubrics, costs a fraction of a physician, and nobody is recruiting them for this.
Reach does not reach 5 because there is no register that identifies a physician as a heavy documenter or an experienced note reviewer. You can verify a licence; you cannot verify the thing that makes someone a good rater here.
Where the benchmarks sit
| Benchmark | Status | Frontier score |
|---|---|---|
| Vals MedScribe (100 rubrics, 87 models) | Near-saturated — leaders within ~3 points | Claude Opus 5 90.99% |
| ACI-Bench (Microsoft/Nuance) | Public and saturated — do not build here | — |
| MedHELM Clinical Note Generation | Strong category; six datasets, four gated or private | Top 0.85 (DeepSeek R1, o3-mini) |
| Vals MedCode — the contrast | Wide open | Claude Opus 5 63.57% |
Sources: MedScribe, MedCode, ACI-Bench, MedHELM.
The residual headroom on MedScribe is concentrated in the Plan section — the part of the note that is clinical reasoning rather than transcription. That is a genuine finding and it is also a redirection: the unsolved part of note generation is really Clinical reasoning wearing a documentation label. See The measurement landscape.
What would kill it
Room scores 2 and defense scores 2. It is already happening.
Saturation. A leaderboard where the top models cluster within three points of 91% does not support a ground-truth product. It supports a disagreement product — harder cases, adversarial encounters, the Plan section specifically — which is smaller, differently priced, and which nobody has been shown buying.
In-house capability. Ambience runs clinical AI as a permanent internal discipline; Abridge employs clinician scientists; Microsoft/Nuance publishes its own dataset. The two documented external purchases were both validation studies for publication, which is exactly the kind of work a company commissions once per product claim.
Physician pricing. Two raters plus an adjudicator at $130–250/hr against a buyer whose alternative is asking three of its clinical advisory board to do it.
The one genuine tailwind is regulatory heat, and it is real. Ontario's Auditor General found an AI medical transcriber for Ontario doctors "hallucinated" and generated errors (CBC, 13 May 2026); Nature published on barriers to scaling ambient AI scribes across diverse settings (23 March 2026); Nurse.org reported that AI charting tools show racial bias and make errors while nurses remain liable (14 April 2026). Every one of those creates budget for third-party note evaluation — and it is the argument for revisiting this page in 2027 rather than entering in 2026. See The regulator wrote your product spec.
Abridge is the largest buyer in this niche at a $5.3B valuation, and whether it uses external clinician panels is unknown — its Greenhouse board 404s on the canonical path, and its own engineering blog describes "structured feedback provided by expert clinical reviewers" without saying whether those reviewers are employees, contractors or vendor-supplied. Nabla and Corti are unresolved in both directions. And no rate is published for note rating anywhere: Suki's 388 paired reviews and GraiLabs' 49 adjudicated encounters both came with no compensation disclosure, so the per-review price in the one niche with documented purchasing is still unknown.
Where the record is thin
The two documented purchases are both single studies tied to a publication. Nothing establishes whether either company commissions rating work on an ongoing basis, at what cadence, or at what volume — and a validation study is a very different business from a standing panel.
Nothing is known about how much of the evaluation labour in this niche is not purchased at all: contributed by health-system partners during procurement bake-offs, or by clinical advisory boards as part of an equity relationship. Any sizing that ignores that free supply overstates the market. Compare The health read and Characterised disagreement.