Miju Labs

All specialists

Vals AI

$40M at $400M for building the benchmark layer in professions the labs were not yet buying for. The existence proof that a specialist can manufacture a budget rather than wait for one.

high confidence4 minupdated 2026-08-30benchmarks · evals · law · finance · clinical
Vertical
Multi-vertical benchmarking
Founded
Not disclosed
Headquarters
Not disclosed
Raised
$40M Series A
Last valuation
$400M
Revenue
Not disclosed in absolute terms — reported up 8x versus all of 2025
Status
Active — cited in model cards from five labs
Who runs it · 2 people in the index

Latest (Sep 2026): On 2026-08-13 Vals announced a $40M Series A at a $400M valuation led by a16z (with 8VC, Pear VC, Bloomberg Beta, HRT Ventures, Next Ladder), after a ~$5M seed in 2025; it reports revenue up 8x vs 2025 and customers doubled. In May 2026 it retired its CorpFin benchmark as saturated and replaced it with an Excel-modelling test. TechCrunch (2026-09-19) reports headcount grew from 8 at the start of 2026 to 25, plans for 10-15 more hires, a new federal-agency model evaluation program, and a move to a larger SF office on Folsom Street. Open roles include Head of Research and Head of Operations.

What Vals AI is saying
Vals AI
10,294 followers
AI models are advancing faster than many of the benchmarks meant to measure them. TechCrunch visited our office to look at how we’re building a more neutral, rigorous and trustworthy evaluation layer for AI. Our work goes beyond abstract tests of intelligence to examine whether models can perform real-world work — and what risks emerge when they are deployed — across myriad industries including law, finance, coding, cybersecurity, biosecurity, and mental health. Thank you Lucas Ropek for spending time with our team and telling our story.
1
AI models are advancing faster than legacy benchmarks can keep up. @TechCrunch @LucasRopek1 visited us to see how we’re building independent evaluations grounded in real-world work and designed to measure both capability and risk
Rayan K.
Evaluating LLMs with Vals AI
Independent evaluation is only credible when evaluators have meaningful access, real independence, transparency about the terms of their work, and the freedom to publish their findings. I signed the AI Evaluator Forum’s statement because these principles are closely aligned with how we think about evaluation at Vals AI. If embedded evaluation is going to strengthen trust in frontier AI, the conditions have to be right.
31 comments

Vals AI raised a $40M Series A at a $400M valuation led by a16z in August 2026, with 8VC, Bloomberg Beta, HRT Ventures and Next Ladder Ventures. Reported alongside it: revenue up 8x versus all of 2025, customer base doubled, team tripled in six months. It benchmarks law, banking, engineering and medicine, works with reference institutions on task taxonomies, keeps private test sets to prevent contamination, and its results appear in model cards from OpenAI, Anthropic, Google, Meta and xAI. It also advises the US Department of Commerce (Pulse 2).

Note the revenue disclosure carefully: 8x growth is a multiple, not an amount. No absolute revenue, gross or net, is public. A $400M valuation on an undisclosed base is a multiple of nothing you can check.

Why this page matters more than its size suggests

Vals did not enter a domain where labs were already spending. It built the benchmark layer for professions the labs were not yet buying for, and the spending followed.

That is the single most useful precedent in the atlas for a vertical specialist, because it inverts the usual sequencing. The default move is to find a niche with proven budget and compete for it — which is Offensive security, where Irregular and Gray Swan AI already own the reference benchmarks, or Investment banking and financial modelling, where OpenAI has hired the subject-matter expert in-house. Vals took the opposite side: pick domains where the benchmark layer is empty, become the reference, and let procurement organise itself around your numbers.

The atlas-wide observation this supports: the benchmark gap and the demand signal are inversely correlated. Where labs spend most, vendors already own the benchmarks. Where the benchmark layer is empty — Law, Accounting, audit and tax, defensive security — it is empty because the budget has not arrived. Whichever niche you pick, you are choosing between fighting for a proven budget and manufacturing one. Vals is the proof that manufacturing works, and it is discussed further in The specialist wedge.

The four verticals, and what each benchmark actually found

Law. The Vals Legal AI Report (VLAIR), October 2025, tested Alexi, Counsel Stack, Midpage and ChatGPT against a US law firm baseline using 200 Q&A sets from consortium firms, graded blind by lawyers and law librarians including evaluators from LegalBenchmarks.ai and the Vanderbilt AI Law Lab, weighted 50% accuracy / 40% authoritativeness / 10% appropriateness. AI products scored 74–78% weighted; the lawyer baseline was 69% (Vals AI).

Finance. The Finance Agent Benchmark v2 covers nine analytical categories including DCF/NPV, LBO and M&A accretion/dilution models, and credits the financial experts who built it (Vals AI). Best model: 60.60% partial credit, with the two hardest categories Financial Modeling at 34.52% and Precedents at 36.37%. Separately, Vals reports frontier models correctly completing fewer than 52% of real financial analysis tasks, and retired its CorpFin benchmark in May 2026 because it "stopped producing meaningful differentiation", replacing it with an Excel-modelling test (AI Business Weekly).

Medicine and engineering are benchmarked but no scores are in the record here.

Two things to take from the VLAIR result specifically. First, the finding is favourable to the vendors being tested — the products beat the human baseline — which is why firms cooperate. A benchmark that only humiliates its subjects gets no participation. Second, VLAIR is private and consortium-run, and is already the reference point for legal AI procurement. That is the position: not a leaderboard people admire, a number people buy against.

What Vals is not

It is a benchmarking company, not a data foundry. It measures; it does not sell the underlying expert supply. In Law, where no law-only expert-data company exists at all, that leaves the data-supply position underneath Vals empty — the benchmark says the models fail at drafting and redlining, and nobody is selling the lawyer-judgement corpus that would fix it.

The retirement of CorpFin is the operational lesson. A benchmark saturates, stops discriminating, and has to be replaced — which means the eval layer is a treadmill, not an annuity. That is manageable at $40M of capital and brutal at seed scale, and it is the reason the atlas repeatedly warns that evals get you the meeting rather than the invoice.

What is not known

Absolute revenue, gross or net — only the 8x multiple. Customer names and count beyond "doubled". Contract values and pricing. Founding date, headquarters and headcount. How Vals compensates the lawyers, bankers, clinicians and engineers who build and grade its private test sets — the entire supply-side economics of the company are undisclosed.

What to learn from it: if you cannot find a niche where the budget already exists, build the number the buyer will be judged against — and make the first published result flattering enough that the industry participates in it.