Centaur Labs
The clinical incumbent: $31M raised to crowdsource medical annotation through a diagnosis game. It sells labelled artefacts, which leaves physician reasoning unsold.
Latest (Sep 2026): No major 2026 funding or leadership event found; the last disclosed round is the $16M Series B (announced Oct 2024), bringing total raised to roughly $31-35M. The site now positions Centaur beyond healthcare (LLMs/software, insurance, robotics, finance) and claims 100,000+ expert annotators and clients such as Microsoft, NIH, Medtronic and Memorial Sloan Kettering. A March 2026 blog post with CEO Erik Duhaime pitches dynamic, disagreement-driven labeling pipelines. No evidence found of named frontier-lab contracts.
Centaur Labs, now trading as Centaur AI, raised a $15M Series A led by Matrix Partners and subsequently a $16M Series B, for crowdsourced medical data annotation delivered through its DiagnosUs app and sold to health and life-science AI developers (Centaur; BioSpace).
The supply mechanism is the interesting part and the most copied. DiagnosUs is a competitive diagnosis app for medical students and clinicians — a game with a leaderboard that happens to emit labelled medical data. It is the Which side you build first answer for a domain where you cannot simply post a rate card, and it is structurally the same trick as Gray Swan's Arena in Offensive security: make the status economy do the recruiting.
What it sells, and the gap that leaves
Centaur sells labelled artefacts — images, scans, records, annotations on them. So does everyone else in the Clinical medicine lane. Protege raised $30M led by a16z for a governed marketplace for AI training data with an explicit healthcare vertical and partnerships with Segmed and Gradient Health on medical imaging (SiliconANGLE). OmicsBank raised $2.25M to expand clinical data infrastructure for healthcare, life sciences and "frontier AI labs" (PR Newswire).
None of them sells physician reasoning traces, diagnostic deliberation, or specialist second opinions as structured expert data. That is an open position, and it is the clinical equivalent of the point Contra Labs misses in design: the labelled artefact is a commodity, the professional's narrated judgement is not.
The buyer is real and unusually explicit
Clinical has the best-documented lab demand of the credentialed professions.
OpenAI's HealthBench was built with 262 physicians who have practised in 60 countries, spanning 49 languages and 26 specialties, producing 5,000 realistic health conversations and 48,562 unique physician-written rubric criteria (OpenAI). OpenAI hires Research Engineer/Scientist, Health AI at $310,000–$460,000 (Becker's). Google DeepMind's AMIE was published in Nature off a randomised, double-blind crossover OSCE with 20 board-certified primary care physicians, 20 patient-actors and 159 case scenarios, graded by majority vote of three specialists across 32 axes (Nature). xAI named medicine as one of four specialist tutor domains (TechCrunch). Mercor's APEX used Mount Sinai Hospital graders (TIME).
Note that both HealthBench and AMIE used synthetic conversations and patient-actors, never real patients. The entire HIPAA problem is avoidable by construction, which is why the niche is open despite having the worst compliance profile in the atlas.
The cost line is the constraint
| Benchmark | Rate |
|---|---|
| BLS physicians and surgeons, median | $132.66/hr ($275,930/yr, 862,800 US jobs) |
| Locum tenens, all-specialty average | $215/hr |
| Locum gastroenterology / anaesthesiology / radiology | $367 / $292 / $289 per hour |
| Locum family medicine / general paediatrics | $140 / $108 per hour |
| Mercor advertised expert ceiling | $200+/hr |
Sources: BLS; Physician Side Gigs, June 2025; Sacra.
A gastroenterologist's locum rate exceeds Mercor's advertised ceiling. That is the structural fact governing clinical supply: you must segment. Primary care, paediatrics and psychiatry clear at data rates; procedural specialties do not, unless you buy per-artefact rather than per-hour. Centaur's game-plus-crowd model sidesteps this by recruiting heavily from students and trainees, which lowers the cost basis and changes what the labels are worth.
The benchmark position
MedQA is saturated — frontier models above 95%. HealthBench is open: GPT-3.5 Turbo 16% → GPT-4o 32% → o3 60%, with OpenAI stating HealthBench Hard leaves "plenty of headroom" (OpenAI). MedHELM (Stanford CRFM and Pacific AI, 40+ clinical scenarios) puts its Q2 2026 leader at 0.652 (Pacific AI). The reasoning layer is unmeasured and unsold.
Revenue, ARR and gross margin — none disclosed at any point. Customer names and count. Valuation at either round. Founding date, headquarters and headcount. What DiagnosUs participants are actually paid, and what share of Centaur's top line reaches them — the take rate on a gamified crowd is undisclosed here and is the number that determines whether this model is replicable.
What to learn from it: in a domain where the wage floor is $132/hr, the way in is not a better rate — it is a status mechanic that recruits trainees, plus a synthetic-vignette construction that removes the compliance blocker entirely.