Miju Labs

All specialists

Centaur Labs

The clinical incumbent: $31M raised to crowdsource medical annotation through a diagnosis game. It sells labelled artefacts, which leaves physician reasoning unsold.

medium confidence4 minupdated 2026-08-30clinical · annotation · medical imaging · crowds
Vertical
Clinical
Founded
Not disclosed
Headquarters
Not disclosed
Raised
~$31M — $15M Series A plus $16M Series B
Last valuation
Not disclosed
Revenue
Not disclosed
Status
Active — trading as Centaur AI
Who runs it · 3 people in the index

Latest (Sep 2026): No major 2026 funding or leadership event found; the last disclosed round is the $16M Series B (announced Oct 2024), bringing total raised to roughly $31-35M. The site now positions Centaur beyond healthcare (LLMs/software, insurance, robotics, finance) and claims 100,000+ expert annotators and clients such as Microsoft, NIH, Medtronic and Memorial Sloan Kettering. A March 2026 blog post with CEO Erik Duhaime pitches dynamic, disagreement-driven labeling pipelines. No evidence found of named frontier-lab contracts.

What Centaur Labs is saying
Tom Gellatly reposted this
Centaur.ai
7,424 followers
We added factual errors to AI health answers. People still preferred them over another answer nearly three times out of ten. In the first two studies of our series, The Limits of Preference, we examine what people’s preferences can tell us and why medical accuracy needs its own evaluation. Read the first study: lnkd.in/eCgsHB-K Read the second study: lnkd.in/eR3wycvC
259 reposts
Tom Gellatly reposted this
Centaur.ai
7,424 followers
Three new datasets just landed in Centaur’s Arena: 🔬 Lymph node tissue slices 🩻 Pediatric chest X-rays 🧠 EEG recordings Across all three, Centaur’s human consensus outperformed every frontier model tested -- with gaps ranging from 7.8 to 35.1 percentage points. Take a look at the Arena and see how today’s leading models stack up against human expertise! Full story in the comments 👇
321 comments11 reposts

Centaur Labs, now trading as Centaur AI, raised a $15M Series A led by Matrix Partners and subsequently a $16M Series B, for crowdsourced medical data annotation delivered through its DiagnosUs app and sold to health and life-science AI developers (Centaur; BioSpace).

The supply mechanism is the interesting part and the most copied. DiagnosUs is a competitive diagnosis app for medical students and clinicians — a game with a leaderboard that happens to emit labelled medical data. It is the Which side you build first answer for a domain where you cannot simply post a rate card, and it is structurally the same trick as Gray Swan's Arena in Offensive security: make the status economy do the recruiting.

What it sells, and the gap that leaves

Centaur sells labelled artefacts — images, scans, records, annotations on them. So does everyone else in the Clinical medicine lane. Protege raised $30M led by a16z for a governed marketplace for AI training data with an explicit healthcare vertical and partnerships with Segmed and Gradient Health on medical imaging (SiliconANGLE). OmicsBank raised $2.25M to expand clinical data infrastructure for healthcare, life sciences and "frontier AI labs" (PR Newswire).

None of them sells physician reasoning traces, diagnostic deliberation, or specialist second opinions as structured expert data. That is an open position, and it is the clinical equivalent of the point Contra Labs misses in design: the labelled artefact is a commodity, the professional's narrated judgement is not.

The buyer is real and unusually explicit

Clinical has the best-documented lab demand of the credentialed professions.

OpenAI's HealthBench was built with 262 physicians who have practised in 60 countries, spanning 49 languages and 26 specialties, producing 5,000 realistic health conversations and 48,562 unique physician-written rubric criteria (OpenAI). OpenAI hires Research Engineer/Scientist, Health AI at $310,000–$460,000 (Becker's). Google DeepMind's AMIE was published in Nature off a randomised, double-blind crossover OSCE with 20 board-certified primary care physicians, 20 patient-actors and 159 case scenarios, graded by majority vote of three specialists across 32 axes (Nature). xAI named medicine as one of four specialist tutor domains (TechCrunch). Mercor's APEX used Mount Sinai Hospital graders (TIME).

Note that both HealthBench and AMIE used synthetic conversations and patient-actors, never real patients. The entire HIPAA problem is avoidable by construction, which is why the niche is open despite having the worst compliance profile in the atlas.

The cost line is the constraint

BenchmarkRate
BLS physicians and surgeons, median$132.66/hr ($275,930/yr, 862,800 US jobs)
Locum tenens, all-specialty average$215/hr
Locum gastroenterology / anaesthesiology / radiology$367 / $292 / $289 per hour
Locum family medicine / general paediatrics$140 / $108 per hour
Mercor advertised expert ceiling$200+/hr

Sources: BLS; Physician Side Gigs, June 2025; Sacra.

A gastroenterologist's locum rate exceeds Mercor's advertised ceiling. That is the structural fact governing clinical supply: you must segment. Primary care, paediatrics and psychiatry clear at data rates; procedural specialties do not, unless you buy per-artefact rather than per-hour. Centaur's game-plus-crowd model sidesteps this by recruiting heavily from students and trainees, which lowers the cost basis and changes what the labels are worth.

The benchmark position

MedQA is saturated — frontier models above 95%. HealthBench is open: GPT-3.5 Turbo 16% → GPT-4o 32% → o3 60%, with OpenAI stating HealthBench Hard leaves "plenty of headroom" (OpenAI). MedHELM (Stanford CRFM and Pacific AI, 40+ clinical scenarios) puts its Q2 2026 leader at 0.652 (Pacific AI). The reasoning layer is unmeasured and unsold.

What is not known

Revenue, ARR and gross margin — none disclosed at any point. Customer names and count. Valuation at either round. Founding date, headquarters and headcount. What DiagnosUs participants are actually paid, and what share of Centaur's top line reaches them — the take rate on a gamified crowd is undisclosed here and is the number that determines whether this model is replicable.

What to learn from it: in a domain where the wage floor is $132/hr, the way in is not a better rate — it is a status mechanic that recruits trainees, plus a synthetic-vignette construction that removes the compliance blocker entirely.