Vals AI
$40M at $400M for building the benchmark layer in professions the labs were not yet buying for. The existence proof that a specialist can manufacture a budget rather than wait for one.
Latest (Sep 2026): On 2026-08-13 Vals announced a $40M Series A at a $400M valuation led by a16z (with 8VC, Pear VC, Bloomberg Beta, HRT Ventures, Next Ladder), after a ~$5M seed in 2025; it reports revenue up 8x vs 2025 and customers doubled. In May 2026 it retired its CorpFin benchmark as saturated and replaced it with an Excel-modelling test. TechCrunch (2026-09-19) reports headcount grew from 8 at the start of 2026 to 25, plans for 10-15 more hires, a new federal-agency model evaluation program, and a move to a larger SF office on Folsom Street. Open roles include Head of Research and Head of Operations.
Vals AI raised a $40M Series A at a $400M valuation led by a16z in August 2026, with 8VC, Bloomberg Beta, HRT Ventures and Next Ladder Ventures. Reported alongside it: revenue up 8x versus all of 2025, customer base doubled, team tripled in six months. It benchmarks law, banking, engineering and medicine, works with reference institutions on task taxonomies, keeps private test sets to prevent contamination, and its results appear in model cards from OpenAI, Anthropic, Google, Meta and xAI. It also advises the US Department of Commerce (Pulse 2).
Note the revenue disclosure carefully: 8x growth is a multiple, not an amount. No absolute revenue, gross or net, is public. A $400M valuation on an undisclosed base is a multiple of nothing you can check.
Why this page matters more than its size suggests
Vals did not enter a domain where labs were already spending. It built the benchmark layer for professions the labs were not yet buying for, and the spending followed.
That is the single most useful precedent in the atlas for a vertical specialist, because it inverts the usual sequencing. The default move is to find a niche with proven budget and compete for it — which is Offensive security, where Irregular and Gray Swan AI already own the reference benchmarks, or Investment banking and financial modelling, where OpenAI has hired the subject-matter expert in-house. Vals took the opposite side: pick domains where the benchmark layer is empty, become the reference, and let procurement organise itself around your numbers.
The atlas-wide observation this supports: the benchmark gap and the demand signal are inversely correlated. Where labs spend most, vendors already own the benchmarks. Where the benchmark layer is empty — Law, Accounting, audit and tax, defensive security — it is empty because the budget has not arrived. Whichever niche you pick, you are choosing between fighting for a proven budget and manufacturing one. Vals is the proof that manufacturing works, and it is discussed further in The specialist wedge.
The four verticals, and what each benchmark actually found
Law. The Vals Legal AI Report (VLAIR), October 2025, tested Alexi, Counsel Stack, Midpage and ChatGPT against a US law firm baseline using 200 Q&A sets from consortium firms, graded blind by lawyers and law librarians including evaluators from LegalBenchmarks.ai and the Vanderbilt AI Law Lab, weighted 50% accuracy / 40% authoritativeness / 10% appropriateness. AI products scored 74–78% weighted; the lawyer baseline was 69% (Vals AI).
Finance. The Finance Agent Benchmark v2 covers nine analytical categories including DCF/NPV, LBO and M&A accretion/dilution models, and credits the financial experts who built it (Vals AI). Best model: 60.60% partial credit, with the two hardest categories Financial Modeling at 34.52% and Precedents at 36.37%. Separately, Vals reports frontier models correctly completing fewer than 52% of real financial analysis tasks, and retired its CorpFin benchmark in May 2026 because it "stopped producing meaningful differentiation", replacing it with an Excel-modelling test (AI Business Weekly).
Medicine and engineering are benchmarked but no scores are in the record here.
Two things to take from the VLAIR result specifically. First, the finding is favourable to the vendors being tested — the products beat the human baseline — which is why firms cooperate. A benchmark that only humiliates its subjects gets no participation. Second, VLAIR is private and consortium-run, and is already the reference point for legal AI procurement. That is the position: not a leaderboard people admire, a number people buy against.
What Vals is not
It is a benchmarking company, not a data foundry. It measures; it does not sell the underlying expert supply. In Law, where no law-only expert-data company exists at all, that leaves the data-supply position underneath Vals empty — the benchmark says the models fail at drafting and redlining, and nobody is selling the lawyer-judgement corpus that would fix it.
The retirement of CorpFin is the operational lesson. A benchmark saturates, stops discriminating, and has to be replaced — which means the eval layer is a treadmill, not an annuity. That is manageable at $40M of capital and brutal at seed scale, and it is the reason the atlas repeatedly warns that evals get you the meeting rather than the invoice.
Absolute revenue, gross or net — only the 8x multiple. Customer names and count beyond "doubled". Contract values and pricing. Founding date, headquarters and headcount. How Vals compensates the lawyers, bankers, clinicians and engineers who build and grade its private test sets — the entire supply-side economics of the company are undisclosed.
What to learn from it: if you cannot find a niche where the budget already exists, build the number the buyer will be judged against — and make the first published result flattering enough that the industry participates in it.
The specialist wedge
The bet is that one domain buys cheaper experts and faster belief, and that both advantages expire the moment you have a reference customer. What would have to be true, what the evidence supports, and the trade-off that decides which niche.
Accounting, audit and tax
The cheapest credentialed pool in the set, the only one with a queryable national licence register, and nobody selling into it — against the second-thinnest evidence that a frontier lab wants it.
Law
The largest credentialed pool anywhere in this atlas, a real observed clearing price of $140–160/hr, no vertical specialist — and a privilege problem with the cleanest workaround in the report.
What you can actually sell
There is no public per-environment price for a cyber range anywhere. Every cyber figure that exists is a contract envelope, an evaluation run cost, or a general-RL price applied by analogy — and the artefacts that price well are the ones no machine can grade.
Irregular
The company you have filed as Pattern Labs. Same firm, new name: 35 people in Tel Aviv running cyber evals for four frontier labs, at $450M.
Publishing the benchmark
Yes, benchmarks convert — but only with a commercial vehicle attached. Vals AI raised at $400M on the explicit logic that public benchmarks make private evals sellable; METR is the most-cited organisation in the field and takes no lab money at all.
Clinical medicine
The best-documented lab demand in the set and the widest open benchmark among the professional domains — attached to the highest cost basis anywhere, where a gastroenterologist's opportunity cost exceeds Mercor's entire ceiling.
Halluminate
Occupies the finance niche by name, with $160K of disclosed funding that is almost certainly stale. Whether investment banking is taken or barely touched turns on a number nobody has.
Edison Scientific
$70M and already contracting PhD biologists to build its own benchmark. The wet-lab position is not open — it is occupied by a company that built the moat before selling anything.
Investment banking and financial modelling
The only domain in this research where a frontier lab keeps permanent in-house subject-matter headcount — and the only one already occupied by a YC company selling the exact same thing under its own name.