Miju Labs

The security dossier

Behavioural and mental health

The only niche where frontier labs buy directly and say so in writing, and the highest willingness-to-pay multiple in the dossier — occupied by a funded vertical specialist with a psychologist CEO, a published benchmark and foundation-model customers.

watchhigh confidence8 minupdated 2026-08-30

OpenAI published the numbers. Its October 2025 sensitive-conversations work drew on a Global Physician Network of "nearly 300 physicians and psychologists who have practiced in 60 countries," of whom over 170 clinicians — psychiatrists, psychologists and primary-care practitioners — wrote ideal responses, analysed model outputs and rated safety, evaluating more than 1,800 model responses in serious mental-health situations. The disclosed prevalence: psychosis or mania in ~0.07% of weekly active users, suicidal ideation with explicit planning or intent in 0.15%, emotional reliance in ~0.15%. The result: 65% fewer non-compliant responses for psychosis/mania, 65% for self-harm and suicide, 80% for emotional reliance (OpenAI).

Anthropic says the same thing in partnership form: "We partner with ThroughLine, a leader in online crisis support, to develop a deep understanding of where and how models should respond in situations related to self-harm and mental health" (Anthropic). Affective conversations are 2.9% of Claude.ai conversations — 131,484 of roughly 4.5 million analysed (Anthropic).

This is the only niche in this dossier where frontier labs are the confirmed, named, direct buyer. Everywhere else the demand is inferred from Mercor's rate card. Here two labs have written it down.

And it is already occupied.

What the artefact is

Five things, all simulated, none of them a real therapy transcript:

  1. Multi-turn crisis-escalation transcripts with a clinician-graded rubric applied per turn, not per conversation.
  2. Risk-assessment labels — suicidality, self-harm, psychosis and mania, eating-disorder cues, emotional reliance — with severity and the correct clinical action attached.
  3. Therapeutic-modality fidelity annotations: whether a response adheres to CBT, ACT or motivational-interviewing technique.
  4. Diagnostic-interview traces.
  5. Preference pairs authored by licensed clinicians for reward modelling.

The per-turn structure is the technically distinctive part. A crisis conversation fails at a specific turn, and the commercially useful label locates the failure rather than scoring the transcript. mpathic's own benchmark is built this way, and so is OpenAI's rating protocol.

Does it need patient data

No, and the "no" here is stronger than anywhere else in the dossier.

Real therapy transcripts are the most legally radioactive data in healthcare — psychotherapy notes carry heightened protection even inside HIPAA, and nobody credible is buying them. Every serious artefact in this niche is simulated, and clinician-authored simulated conversations are the established standard of practice rather than a workaround. mPACT, CounselBench and OpenAI's 1,800 rated responses are all built this way.

The affirmative argument matters here more than anywhere. Constructed conversations permit controlled variation no naturally occurring corpus allows: matched pairs identical but for one attribute, escalation probes at a chosen turn, difficulty calibration, contamination-free held-out sets. You cannot randomise a real patient's presentation while holding everything else fixed. See Build it without ever touching a patient record.

What replaces the PHI problem is a human-subjects-adjacent problem of a different kind: the annotators themselves. Vicarious trauma is a known occupational hazard of exactly this work, and an annotator reading crisis material for eight hours is a person you owe a duty of care to. That is an operating cost, and it is the reason cost scores 3 rather than 5 despite the labour being cheap.

Is anyone buying

Budget scores 5 — the only 5 in this set of seven, and one of two in the whole dossier alongside Clinical reasoning.

Beyond OpenAI and Anthropic, the application layer is funded and visibly buying clinical input. Slingshot AI (Ash) raised $93M and trained with "a reward model trained on thousands of comparisons written by clinical experts," with a fine-tuning stage in which "clinicians then adjust the model" (Nebius, Behavioral Health Business) — though its own safety study has drawn scepticism (STAT, 24 Nov 2025). Behind it: Jimini Health ($17M, March 2026), The Path ($14.3M, May 2026), Theris (out of stealth, May 2026) and Talkspace's AI therapist (June 2026).

Mercor prices the labour at $80–150/hr for a Child & Adolescent Mental Health Clinical Advisor (Mercor) — against a market wage of $28.53–47.65/hr, the highest willingness-to-pay multiple in the dossier, roughly 2 to 5 times.

Proof scores 3, not higher, and the reason is the next section. There is a public benchmark in this space already, and it is not yours.

What the expert costs

Cost scores 3: cheap people, expensive governance.

BLS, May 2025: substance-abuse, behavioural-disorder and mental-health counsellors median $59,350/yr = $28.53/hr, 533,400 jobs, +18% to 2035 (BLS); psychologists median $99,110/yr = $47.65/hr, 209,400 jobs, with clinical and counselling psychologists at 81,300 jobs and a $100,580 median (BLS).

The labour arbitrage is therefore excellent. What drags the score is everything you must build around it:

  • A named licensed clinical director. Anyone pitching this niche without one should not win the business, and buyers in this space will ask.
  • An escalation protocol for annotators who encounter material they judge to require action, and supervision structure behind it.
  • Annotator mental-health support, budgeted rather than gestured at.
  • An IRB-grade consent posture on anything touching human subjects.
  • Fifty-state licence verification. There is no Nursys equivalent for LPC, LMFT, LCSW or psychologist — verification is state-by-state, with ASPPB's PSYPACT and psychology licensure databank covering psychologists only. That is materially harder and more expensive per candidate than Nursing or Medical coding.

Getting to them

Reach scores 4, one below nursing and coding, and the missing point is the registry.

The channels are good: APA and ACA member directories, NASW for social workers, the state psychological and counselling associations, and the specialty communities around crisis work. PSYPACT gives a multi-state structure for psychologists specifically.

But verification is fifty separate lookups across four professions, and the professions do not share a credential vocabulary. Building the verification pipeline is a real week of work rather than an afternoon, and it has to be re-run per state as licences lapse.

The professional-body posture is workable. The APA issued a health advisory on generative-AI chatbots and wellness apps for mental health on 13 November 2025 — cautionary, not prohibitive — and it runs an AI advisory on which mpathic's Chief Science Officer sits. The APA is engaging, not boycotting. Contrast National Nurses United in Nursing.

Where the benchmarks sit

BenchmarkOwnerAccessNote
mPACTmpathicProprietaryClinician-designed multi-turn conversations across suicide risk, eating disorders and misinformation; each response scored by trained clinicians on a rubric
CounselBenchAcademicPublicDeveloped with 100 mental-health professionals; 2,000 expert evaluations; 1,080 responses from nine LLMs in the adversarial split (arXiv 2506.08584)
PsychBenchAcademicPublicPsychiatric clinical practice (arXiv 2503.01903)
PsychiatryBenchAcademicPublicMulti-task psychiatry, npj Digital Medicine (Nature)
Mindbench.aiAcademicPlatformNature NPP—Digital Psychiatry (Nature)
MedHELM MentalHealthStanfordPrivate, newCounsellor-patient conversations requiring empathetic responses

mPACT's published finding is the one to internalise: tested against Claude Sonnet 4.5, GPT-5.2, Gemini 2.5 Flash, Grok 4.1 and Mistral Medium 3, models avoid overt harm but "consistently fell short of what a clinician would consider an adequate response in a real crisis situation," with Claude Sonnet 4.5 leading on overall clinical alignment (GeekWire). That is precisely the finding a new entrant would have wanted to publish first.

The underlying data landscape is genuinely thin: a survey of 89 clinical mental-health datasets finds most have fewer than 200 participants, with heavy English and Chinese concentration, no standardised annotation protocols, and most datasets private (arXiv 2508.09809). See The measurement landscape.

What would kill it

Room scores 2, and mpathic is why.

Seattle-based mpathic sells exactly the product under evaluation — "clinical-grade evaluation and training datasets for AI safety," human-in-the-loop infrastructure combining synthetic psychological scenarios with licensed-clinician review. $15M raised in 2025 led by Foundry VC, reported 5x quarter-over-quarter growth by end-2025, ~34 employees and "a global network of thousands of licensed clinical experts." CEO Grin Lord is a board-certified psychologist; the Chief Science Officer sits on the APA's AI advisory. Its customers are "foundational AI model developers and LLM-powered application teams serving tens of millions of users," and one early engagement produced a ">70% reduction in undesired or dangerous model responses" (GeekWire).

A funded vertical specialist with a clinician-founder, a published benchmark and foundation-model customers caps room at 2 however good the rest of the page looks — the same arithmetic that caps Clinical reasoning.

The second kill is a safety failure of your own making. This is not a compliance line. You are producing artefacts about suicidality and psychosis; a bad label is a plausible contributor to a death, and a public incident ends the company rather than costing it a customer. The mitigation is the governance stack above, staffed before the first project rather than after it.

Regulation, by contrast, is a tailwind. Illinois' Wellness and Oversight for Psychological Resources Act took effect 1 August 2025: unlicensed entities may not provide AI therapy, and licensed professionals may not use AI to make independent therapeutic decisions, communicate therapeutically, generate treatment recommendations without professional review, or detect emotions or mental states — with civil penalties up to $10,000 per violation (Holland & Knight). New York's AI-companion law, effective 5 November 2025, requires operators to detect and address suicidal ideation; SB 7263 would add a private right of action (Holland & Knight).

Read the direction carefully. These laws restrict AI delivering therapy. They do not restrict clinicians being paid to evaluate AI. Regulation here creates evaluation budget and simultaneously constrains the application companies competing for the same clinician time. Defense scores 4 on the strength of that plus the governance barrier to entry. See Europe prefers the design you were going to build for the European position.

State law is mapped incompletely, and no lab posting names a therapist

Nevada and Utah are reported to have comparable AI-therapy guardrails but statute-level detail was not retrieved. Worse, NCSL's 2025 AI legislation tracker lists no enacted mental-health-AI restriction, which directly contradicts the Illinois enactment above — treat the tracker as incomplete rather than authoritative, and commission a proper fifty-state survey before pricing regulatory work. Separately: OpenAI and Anthropic describe their clinician programmes but name no vendor and publish no rate, and no first-party lab posting names a licensed therapist. As with the RN and the CPC, the price signal is Mercor's alone.

Where the record is thin

Neither lab says whether it paid. OpenAI's HealthBench physician cohort is described as compensated in a later paper; the 170+ clinicians in the sensitive-conversations work are not characterised either way. ThroughLine's relationship with Anthropic is a partnership with no value attached.

mpathic's revenue is unknown — "5x quarter-over-quarter growth" from an undisclosed base is a press figure, not a number, and its customers are characterised rather than named. And the unit economics are missing everywhere: no published price per graded turn, per crisis transcript or per preference pair. Compare The health read and Health-AI companies as buyers.