Hippocratic AI ran the largest documented clinician-evaluation programme in health AI: 6,234 US-licensed clinicians — 5,969 nurses and 265 physicians, averaging 11.5 years of clinical experience, evaluating 307,038 unique calls across four system iterations. Correct-medical-advice rates moved from roughly 80.0% pre-Polaris to 96.79%, 98.75% and finally 99.38%; severe-harm concerns fell from 0.06% to 0.00% (Hippocratic AI). Its safety page puts the cumulative figure higher still: "over 7.7K U.S. licensed clinicians to make 775K test calls" (Hippocratic AI).
That is the demand proof and the competitive problem in one sentence. Six thousand nurses is not a vendor relationship. It is a network Hippocratic built itself, and it monetises the result as a product differentiator.
The second fact is the opening. There is no nursing benchmark. MedHELM's 121-task taxonomy has no nursing category at all — care plans are folded into Clinical Note Generation, triage into Clinical Decision Support (arXiv 2505.23802). The only thing that claims the title is NurseLLM (October 2025), which describes itself as "the first large scale nursing MCQ dataset" and was built by "a multi-stage data generation pipeline" — synthetic, not nurse-authored (arXiv 2510.07173). An entire profession of 3.47 million people has one machine-generated multiple-choice set standing in for it.
What the artefact is
Five things, all inside the nursing scope of practice and all writable at a keyboard:
- Triage disposition labels with escalation rationale — the disposition, the red flag that drove it, and the protocol line that applies.
- Care-plan generation and critique against NANDA-I or a standardised care-plan structure.
- Patient-education rewriting at a specified reading level, with teach-back checks.
- Escalation-judgement traces — "what would make you call the physician now". This is the artefact with no substitute. It is the tacit clinical timing judgement that neither a chart nor a guideline document contains.
- Preference pairs on patient-facing conversational turns, for reward modelling.
Scope discipline is the design constraint. Nurses cannot lawfully produce diagnostic judgement, so every artefact must sit inside triage, care planning, education and escalation. Cross that line and the labels are professionally indefensible regardless of how good they are.
Does it need patient data
No. This is the second-cleanest PHI position in the dossier after Medical coding.
Triage vignettes, care plans and patient-education material are routinely written de novo — writing them is part of nursing education and quality work. A fictional case relates to no individual, so it is not individually identifiable health information; no living individual is a human subject under the Common Rule; and no institutional record was opened, so no employer owns the output. See Build it without ever touching a patient record.
The stronger version of the argument is affirmative rather than defensive. Vignettes are a validated instrument, not a compromise. Peabody et al. compared clinical vignettes, chart abstraction and unannounced standardised patients across four conditions at two VA medical centres and found vignettes scored 71.0% against the standardised-patient gold standard of 76.2%, with chart abstraction trailing at 65.6% — the written hypothetical case beat the real medical record as a measure of actual clinical behaviour (JAMA 2000;283:1715-22). That result is twenty-six years old and it is the citation this business rests on.
Employer IP is the residual risk, as everywhere: nurses must not bring charts or employer protocols. It is a straightforward onboarding control.
Is anyone buying
Budget scores 3, and the 3 is honest rather than grudging.
Hippocratic AI is the largest documented buyer of nursing judgement anywhere, at $126M Series C, a $3.5B valuation and $404M raised, claiming 115M+ clinical patient interactions and 50+ health-system, payor and pharma customers (Business Wire). Its Polaris architecture paper describes registered nurses annotating transcripts, providing preference feedback for RLHF, and performing data review and rewriting — the full expert-data stack, in-house (arXiv 2403.13313).
Frontier labs buy through the intermediaries. Mercor publishes live bands: Nursing Talent Network $60–120/hr, Nurse Practitioner $70–110/hr, Adult Inpatient Nurses (RN) $55–65/hr, Utilisation/Case Management leader $100–150/hr (Mercor). Mercor does not list roles it is not filling, so the demand is real — but this is a rate card, not a lab contract. See The labs as health buyers.
Budget does not reach 4 because the buyer set is one company that built rather than bought, plus an intermediary price sheet. Compare Behavioural and mental health, where two named frontier labs describe their clinician programmes in their own publications, and Medical coding, where a dozen funded application companies are visibly in market.
Proof scores 4. With no incumbent benchmark and no vendor holding the position, a nurse-authored triage-and-escalation set published with inter-rater statistics is the fastest route to being visibly the best source in a domain in this whole dossier. The constraint on proof is not building it; it is finding the reference customer who will say so out loud while the largest buyer runs its own panel.
What the expert costs
Cost scores 5. No equipment, no imaging pipeline, no institutional agreement, no PHI insurance.
BLS, Registered Nurses, May 2025: median $97,550/yr = $46.90/hr; 3,465,400 jobs; +6% 2025–2035, 194,700 new jobs (BLS). Nurse practitioners: median $132,300/yr (~$63.61/hr); 336,300 jobs; +41% growth (BLS).
Against $55–120/hr observed, the multiple runs roughly 1.2x at the inpatient-RN floor to 2.6x at the network ceiling. That is materially worse than coding's 3x and materially better than Pharmacy's 1.1–1.8x. And unlike the physician niches, the arbitrage runs the right way at every point on the band — see What a clinician hour costs.
The practical read: buy at the $55–65/hr inpatient band for volume triage and care-plan work, reserve the $100–150/hr utilisation-management band for adjudication and rubric authorship, and the blended cost lands near $70/hr for a product priced against physician-grade alternatives.
Getting to them
Reach scores 5. Recruiting reach here is trivially easy; recruiting sentiment is the hard part.
Verification is cheap and national. Nursys is NCSBN's national nurse licensure and discipline database with QuickConfirm verification and e-Notify (Nursys, bot-blocked to automated fetch; see NCSBN), and every state board runs an independent public lookup on top of it. A candidate's licence, state and discipline history cost pennies to confirm.
Channels: allnurses.com, r/nursing (very large, very vocal), Nurse.org, the state nurses associations, Sigma Theta Tau, and the specialty bodies — ENA for emergency, AACN for critical care, AWHONN for obstetric — plus the enormous specialty Facebook groups.
The message has to be right the first time. National Nurses United's position is that AI is being used to displace nursing judgement; the pitch that works is "you are the oversight", never "help us automate you." Get that wrong publicly once and allnurses and r/nursing close together.
Where the benchmarks sit
| Benchmark | Status | Note |
|---|---|---|
| NurseLLM (Oct 2025) | The only nursing benchmark | Synthetic MCQ, machine-generated pipeline (arXiv 2510.07173) |
| MedHELM | No nursing category exists | Care plans folded into note generation; triage into decision support |
| ED triage studies | Physician-framed, not nursing-framed | 39,375-patient ED LLM triage study; specialty triage competition |
| Schmitt-Thompson, ClearTriage | Proprietary licensed content | A licensing route, not a free one (ClearTriage) |
An arXiv sweep for nursing benchmarks in 2026 returns scheduling work, maternal/neonatal RAG, a Zanzibar on-device RAG deployment and a Japanese multi-profession licensing set. No US nursing-practice benchmark exists. See The measurement landscape.
Note the last row carefully. The established nurse-line triage protocols are commercial products. A nurse-authored escalation benchmark must be built from clinical judgement rather than transcribed from Schmitt-Thompson, and the annotation protocol has to make that distinction explicit — the same discipline Pharmacy needs around Lexicomp.
What would kill it
Defense scores 3 and room scores 4.
The union is the defining risk. National Nurses United surveyed 2,300+ RNs between 18 January and 4 March 2024: 60% disagreed that employers prioritise patient safety in AI implementation; 69% of nurses using algorithmic acuity systems said their own assessments do not match the computer-generated ones; 48% reported mismatches with automated handoffs; 29% cannot modify AI-generated assessments; 40% cannot adjust AI-generated outcome-prediction scores. NNU calls for "an immediate pause on implementing A.I. in health care settings" and a Nurses' and Patients' Bill of Rights (NNU, NNU AI page).
Read it precisely, because the read decides the business. This is opposition to AI replacing nurses at the bedside. It is not opposition to nurses being paid to evaluate AI — and the liability argument runs the same direction: "AI Charting Tools Show Racial Bias and Make Errors—Nurses Are Still Liable" (Nurse.org, 14 April 2026). Nurses already carry the liability. That is the argument for paying them to audit. But a single badly worded recruitment campaign converts a tailwind into a boycott. See What the doctor on the other end is risking.
Defense is capped at 3 by pool size. Three and a half million RNs with free national verification is the opposite of scarce supply. What sustains the position is that escalation judgement is tacit and un-corpused — there is nowhere to scrape it from — not that the people are hard to find.
The concentration risk is Hippocratic. The largest buyer has vertically integrated the function. Sell to the labs and to the second tier of patient-facing health-AI companies; do not build a plan around selling nursing evaluation to the company that has 6,000 nurses of its own.
Neither Hippocratic's RWE-LLM page nor the Polaris paper states whether the 6,234 clinicians were compensated, or at what rate. This is the single most decision-relevant unknown in the niche. If they are paid, it establishes the market rate for nursing evaluation at scale. If they are not — if 307,038 calls were reviewed for access, credit and professional interest — it establishes that clinicians will do this work unpaid, which destroys the pricing model outright. At 6,234 people and 307k calls, unpaid participation is implausible, but implausible is not disproved. And separately: no first-party lab posting anywhere names an RN. The lab demand rests entirely on Mercor's rate card.
Where the record is thin
The national active-RN licence count is unretrievable. NCSBN publishes Active RN Licenses statistics only as downloadable PDF/XLS behind a path that does not resolve to automated fetch (NCSBN), so the BLS employment figure of 3,465,400 is standing in for a licence count that is certainly larger — Veterinary medicine shows the gap can run 45% between licensed and employed.
No price exists per triage vignette, per adjudicated escalation or per care plan. Mercor's hourly bands are the only price signal, and hourly bands do not tell you what a 2,000-case benchmark costs to build or what it sells for.
And nothing is published on how nursing artefacts perform as training signal rather than as evaluation. Every argument here is about building an evaluation set because that is what the evidence supports. Compare the ranked read in The health read and the entry sequence in Characterised disagreement.