Centaur.ai — formerly Centaur Labs — is the company most people name when asked who already does medical data for AI. It is a real business with real customers and eight years of operating history. It is also, on the evidence, not competing for the same dollar.
The decisive fact is what its supply side is and what that supply can produce. DiagnosUs, Centaur's mobile app, gamifies annotation into "contests" where "users across the world compete against each other for cash prizes" (Centaur docs). About half of active participants are US-based; the population skews to medical students, residents and healthcare professionals — but the app is open to anyone and requires no medical degree (docs). Every labeller is continuously scored against seeded gold-standard cases, and only high performers' opinions are weighted in the aggregate. Payment is pay-for-performance, not pay-per-label (SignalFire).
That mechanism is elegant, and it is a mechanism for extracting a reliable majority signal from a semi-credentialed crowd. It is not a mechanism for eliciting a board-certified specialist's reasoning, and the price tells you so.
Centaur prices volume label throughput from a semi-credentialed crowd. Mercor prices a licensed physician's reasoning at $110–250/hr (Mercor). Those are different goods with different buyers. Centaur is not a competitor to a Layer 3 clinical-judgement business — it is the layer below it, and its cost structure does not obviously scale upward into the one above.
What DiagnosUs pays, and why that number is soft
This is the weakest-sourced part of the picture and it should be treated accordingly.
- Third-party marketing claims "up to $50 a day just by playing a couple of hours," while noting "there's no guaranteed income — you have to earn it by ranking high in accuracy," with PayPal withdrawal (phmillennia). This is an affiliate-style blog
[WEAK]. - An App Store reviewer reports "$200+ in the last week or so" (App Store). A single self-reported review
[UNVERIFIED]. - Centaur publishes no per-label rate anywhere. Its own documentation covers the labelling mechanism in detail and explicitly omits payment structures, contest payouts, scoring and eligibility.
The honest read is an implied ~$10–25/hour equivalent for a strong performer — one to two orders of magnitude below the physician rate. That gap is the entire strategic story of this space, and it is the arbitrage Centaur's model exists to exploit: buy cheap semi-credentialed attention, aggregate it into a reliable label, sell the label.
Neither of the two earnings sources above should be quoted to a third party. They are directionally consistent with each other and with a gamified pay-for-performance model, and that is all they are.
What it has raised, and from whom
$15M Series A led by Matrix Partners, announced 3 September 2021, with Samsung Next participating (Forbes). $16M+ Series B led by SignalFire, announced 8 October 2024, oversubscribed, with Matrix, Accel, Susa, Omega, Y Combinator, Samsung Next and Alumni Ventures (Series B). Disclosed total ~$31M.
Valuation, revenue and headcount are not disclosed and could not be established from any source worth standing behind.
One detail in the cap table is worth pausing on: angel backers include Alexandr Wang of Scale AI and Tom Lee of One Medical (about). Scale's founder had a read on this vertical between 2021 and 2024 and chose to back rather than build. That is a data point about how the generalist labelling world sizes medical annotation — interesting enough to hold in mind next to Scale AI and One customer is a binary event.
Founder Erik Duhaime's MIT Center for Collective Intelligence PhD work on combining medical opinions with algorithms is the intellectual origin of the company; the name comes from centaur chess. The advisory board includes Thomas Malone, Matthew Lungren MD (Microsoft's Chief Medical Information Officer) and Tina Kapur PhD (about). HQ Boston; Y Combinator alumnus.
Has it moved into evaluation? Yes — but sideways
This is the question that decides whether Centaur is a rival or a neighbour, and the answer is genuinely partial.
The October 2024 Series B already framed the product across "training, evaluation, monitoring and feedback," and pitched on-demand labelling at "model monitoring and evaluation" use cases. The current site carries a dedicated LLM evaluation page, offers benchmarking of a customer's medical model against expert annotations, and runs an interactive "quality playground" (centaur.ai). Network claims moved from "over 50,000 experts" in October 2024 to "100,000+ subject matter experts" in 2026.
So the move is real. But its shape is "benchmark your medical model against expert labels" — a service sold to a company that already has a model and wants a scorecard. It is not "build us a rubric-graded reasoning benchmark" sold to a frontier lab that wants a new instrument. The first is labelling with an evaluation label on it. The second is what The measurement landscape and What actually gets sold describe, and it is a different product with a different buyer and a different margin.
That distinction is the gap a new entrant would be walking into. It is also the reason Centaur's customer list looks the way it does.
Customers: no frontier lab
Named on the Series B and the current site: Microsoft, NIH, Medtronic, Mass General Brigham, Memorial Sloan Kettering, Paige, SciBite (Elsevier), Volastra, Eight Sleep, Activ Surgical, Major League Baseball (Series B, centaur.ai).
Microsoft is the closest thing to a frontier-lab customer, and it is not disclosed as a frontier-lab relationship in any sense. No OpenAI, Anthropic, Google DeepMind or Meta relationship appears anywhere.
That is not a Centaur-specific failure. It is the sharpest finding in the whole competitive map: only Mercor has verified named frontier-lab customers in health. Protege, BioStack, Vals and Sepal all claim lab relationships without naming one. The health vertical has not yet produced a company that can name a lab — see BioStack and Sepal.
Published results are respectable and firmly in Layer 2: Eight Sleep's snore-detection model moved 70% → 93% accuracy; Paige achieved 10× faster annotation and lifted a pathology algorithm's F1 from 0.6 to 0.78 (SignalFire). On-demand labelling delivers "20–30 annotator reads in under 30 minutes" (Series B).
Revenue, valuation, headcount and per-label rate. The only earnings evidence in existence is an affiliate blog and one App Store review, neither of which belongs in a client-facing document. Whether the LLM-evaluation line is a meaningful revenue share or a positioning page cannot be determined from outside.
A well-run Layer 2 business with a clever supply mechanism, a good customer list and no frontier lab on it. It is a competitor for medical-device and health-system annotation budgets, and a non-competitor for rubric-graded clinical reasoning sold to labs. Watch it for one thing only: whether it uses the 100,000-expert claim to move upmarket into judgement. If it does, its cost structure has to change first, and that is visible from outside long before the revenue is.