25-year-old Stanford CS graduate who co-founded Vals, now the most-cited independent benchmark shop for professional-domain AI (law, finance, medicine), with results in lab model cards. He is the face of the 'independent scorekeeper' thesis that won a16z's $40M Series A in August 2026.
high confidenceSan FranciscoAt Vals AI since 2024Updated 2026-09-19
Background
Krishnan studied computer science at Stanford, worked in Stanford's AI lab and at Microsoft as an undergraduate, and interned at Palantir (TechCrunch, Tech Funding News). Per TechCrunch, Vals grew out of his observation that academic benchmarks were not keeping pace with model progress.
Vals' early reputation came from law: its first legal AI benchmark (Feb 2025), run with law firms such as Reed Smith and Fisher Phillips and a lawyer baseline, which Krishnan described as a 'diplomatic mission' to a market full of mistrust. He also hosts Vals' podcast 'The Bench' (per a Vals LinkedIn post).
What they run now
Scaling Vals from 25 people after the Series A: new domain benchmarks (law, finance, healthcare, coding, mental health, cybersecurity, biosecurity, law of armed conflict, recursive self-improvement), the Vals Index, custom coding benchmarks (Vals Smith), and a federal-agency evaluation program. Revenue model is companies paying Vals to evaluate their models, with test sets kept private.
Career
2024 – presentCo-founder & CEO, Vals AI
—Intern, Palantir
—Undergraduate role, Microsoft
—Undergraduate researcher, Stanford AI lab
—Stanford University, Computer Science
On the record
Independent evaluation · 2026-08
Every lab claims the smartest model; the market needs a neutral scorekeeper (a16z compares the role to Moody's/S&P). source
Real-world impact · 2026-09
Benchmarks should measure the real impact of models on actual work, not academic puzzles. source
Private test sets · 2026-09
Tests are kept undisclosed so they cannot be gamed like public legacy benchmarks. source
Legal AI trust · 2025-02
Law firms mistrust AI tools and are confused about hallucinations; benchmarking with firms builds trust. source
Recent posts
29 posts archived · most engaged first, then the latest
While evaluating Fable 5.1, our team elicited a solution to a cipher that had been open for 373 years, listed among the top 50 unsolved encrypted messages.
It managed this in just 44 minutes and 176k tokens.
Two things stand out to me, and neither is the solve itself.
1) We never pointed the model at this cipher. We asked it, open-endedly, to find and solve an unsolved cipher, and it chose this one. Knowing which problem to work on is one of the biggest bottlenecks in research ability, and Fable showed real intuition for it.
2) The solution itself was quite simple. Many open problems aren't hard, they've just never had enough sustained attention on them. That bottleneck is disappearing, and I'd expect a wave of these long-standing problems to fall quickly.
Our early experiments are also showing Fable 5.1 to have the highest propensity for RSI. More on this to come soon.
Independent evaluation is only credible when evaluators have meaningful access, real independence, transparency about the terms of their work, and the freedom to publish their findings.
I signed the AI Evaluator Forum’s statement because these principles are closely aligned with how we think about evaluation at Vals AI. If embedded evaluation is going to strengthen trust in frontier AI, the conditions have to be right.
Independent evaluation requires meaningful access, real independence, transparency, and the freedom to publish findings. I signed the AEF statement because these principles closely align with how we approach evaluation at Vals.
Jennifer Li — a16z GP, co-author of the Series A announcement
Katie Kirsch — a16z, announced the round on X
Tara Waters — consultant on Vals' legal AI benchmark
Who they amplify
Accounts whose posts Rayan has reposted recently: Vals AI.
Why it matters here
Vals is the proof that a specialist can build the reference benchmark in a domain before labs budget for it, which is exactly the move open to Julian in design/visual taste, where no dominant professional-grade benchmark exists. Krishnan decides which domains Vals enters next; a design/creative benchmark is not on its published list, which is either a gap to own or a partnership opportunity (Vals needs domain experts to build and grade private sets).
How to reach
Speculative: active on X (@RayanKrishnan) and hosts 'The Bench', so a sharp, data-backed proposal for a creative-judgement benchmark (with a panel of credentialed designers) is the kind of thing he might respond to. a16z's infra team is a possible warm path.
What we could not establish
Exact graduation year and degree level at Stanford not verified.
Date of the a16z policy episode not verified (page returned 429).
How Vals recruits and pays its domain experts is undisclosed.