Background
Did his undergraduate degree in electrical engineering at Stanford, then a PhD in EECS at UC Berkeley under Michael I. Jordan and Jitendra Malik, finishing in October 2024, followed by a postdoc with Ion Stoica. His academic reputation rests on statistically rigorous uncertainty quantification: he co-wrote Conformal Prediction: A Gentle Introduction and Prediction-Powered Inference (Science, 2023), and has a Cambridge University Press book on conformal prediction forthcoming in 2026.
Stoica introduced him to Wei-Lin Chiang's Chatbot Arena project; he brought the statistical methodology (Bradley-Terry style rankings, style control) and later the business. Felicis reports he would otherwise have become a professor. He was a student researcher at Google DeepMind in summer 2024.
What they run now
Scaling Arena's commercial evaluation business (paid AI Evaluations, launched September 2025) while defending the neutrality of the public leaderboard, which he frames as un-buyable ('models can't pay to get on, can't pay to get off'). Pushing the platform into new modalities (video, agents, harnesses/tools) and into occupation-specific arenas, and defining 'utility' through personalised, granular evaluations.
Career
- 2025 – presentCo-founder & CEO, LMArena / Arena — Chatbot Arena research project from 2023; company incorporated April 2025
- 2024 – 2025Postdoctoral Scholar (with Ion Stoica), UC Berkeley
- 2024 – 2024Student Researcher, Google DeepMind — summer
- —Stanford University, Undergraduate degree in Electrical Engineering
- completed 2024UC Berkeley, PhD, EECS — advised by Michael I. Jordan and Jitendra Malik
On the record
The public leaderboard is run like public infrastructure; labs cannot pay for placement, and commercial revenue comes from separate evaluation products. source
Argued the Cohere-led 'Leaderboard Illusion' paper contained factual errors and misrepresented open vs. closed model sampling; defends pre-release testing as a feature users like. source
Calls data a 'scaling complement' to compute, says frontier labs spend 10-20% of GPU budgets on data, and puts the market at ~$100B by 2030, possibly far larger; names Mercor, Handshake, Surge and Scale as $1B+ players. source
Arena scores measure utility to its own user base, not abstract quality; style and verbosity are controlled for, with factuality as a separate signal; the hardest open problem is defining utility and personalising evaluations. source
Accepts that labs climb the leaderboard, arguing that if it makes models more useful to Arena's diverse users that is good for the world. source
Says many people still see Arena as an open-source project and do not realise it makes money. source
