Miju Labs

← All people · Snorkel AI

Vincent Sunn Chen

Founding team / co-founder; Research Fellow, leads Open Benchmarks Grants · Snorkel AI

Founding team member (described by Fortune as a co-founder) who leads Snorkel's evaluation and expert-feedback systems and directs the $3M Open Benchmarks Grants. He is Snorkel's public face on benchmark integrity, quoted in Fortune (Sep 2026) on OpenAI changing eval metrics around a launch.

high confidenceAt Snorkel AI since 2019Updated 2026-09-19

Background

Graduate work at the Stanford AI Lab under Chris Ré on multi-task learning, weak supervision and long-tail evaluation; earlier a data engineer on Tesla Autopilot; co-director of Stanford's TreeHacks hackathon.

What they run now

Runs the Open Benchmarks Grants (100+ applications; first wave Jul 2026) and co-authors agent benchmarks (OSWorld 2.0, Agents' Last Exam); hosts the 'Benchtalks' series.

Career

  1. 2019 – presentFounding team; Research Fellow; director, Open Benchmarks Grants, Snorkel AI
  2. Data engineer, Autopilot, Tesla
  3. Stanford UniversityStanford AI Lab, advised by Chris Ré

On the record

Benchmarks · 2026-04

Benchmarks should shape the frontier, not just measure it: they need expert-validated tasks, diverse taxonomies and must stay unsaturated. source

Eval transparency · 2026-09

Calls for clearer industry norms on disclosing post-launch changes to reported benchmark metrics. source

Recent posts

31 posts archived · most engaged first, then the latest

Henry Ehrenberg reposted this
Vincent Sunn Chen
Research & Founding Team @ Snorkel AI
Most benchmarks measure how capable a model is at a fixed distribution of tasks, but capability and improvement are different axes. We're very excited to introduce Continual Learning Bench 1.0 to measure improvement w/ Snorkel AI and University of California, Berkeley! In this benchmark, we introduce novel, expert-validated tasks across real-world domains, where systems are expected to adapt across instances. For instance, a frontier model can be excellent at SWE tasks and still not get any better at your codebase the more it works in it. Work led by Parth Asawa, with Chris Glaze Ramya Ramakrishnan Gabriel Orlanski Benji Xu Asim Biswal Frederic Sala Matei Zaharia Joseph Gonzalez. Learn more at lnkd.in/eJpU-6rv - Accepting more contributions!
1099 comments23 reposts
capability != learning new benchtalks with @pgasawa on continual learning, where we discuss teaching models to learn from experience, measuring learning ability, the bet on parametric models, and more 01:06 What is continual learning? 04:10 Why capability and learning are different 06:13 Why build a benchmark? 08:07 Continual Learning Bench launch and reception 09:13 Anthropic's Fable release and Continual Learning Bench 11:02 How to design tasks for continual learning 18:41 The gain metric 24:01 What good looks like on the leaderboard 29:13 Failure modes: why models can't update their beliefs 31:12 Parametric systems and future architectures 34:30 Open science and AI safety 45:42 Lightning round 49:08 How to contribute to Continual Learning Bench
32210541KView on X ↗
Henry Ehrenberg reposted this
Vincent Sunn Chen
Research & Founding Team @ Snorkel AI
New Benchtalks with John Yang: on ProgramBench (0% frontier models at launch) and the lineage/future of coding benchmarks, from SWE-bench/InterCode to now Full link here: lnkd.in/gg_A8k9j Kudos to the ProgramBench team! + Kilian Lieret Ph.D. (co-lead) Jeffrey Ma Parth Thakkar Dmitrii Pedchenko Sten Sootla Emily McMilin Pengcheng Yin Rui Hou Gabriel Synnaeve Diyi Yang Ofir Press
1055 comments37 reposts
emphasizing the open benchmark QA approach here: - continuous maintenance - all the 'flagged' tasks were all self-reported and triaged by the TB team in the public issue tracker - open traces - everything is on GitHub, so you can audit the feedback, fixes, and back-and-forth. it's 'peer-review' for benchmark tasks! - "trust" needs more nuance - this goes beyond a 'verified/flawed' label. in reality, the reported issues affect <3% of rollouts (this analysis is also public/inspectable) more of this in the open!
prediction: more labs will self-identify as data research labs the overton window has shifted - data was perceived as "plumbing," but now it's sayable (including at frontier labs) that the data is the research!
511764.2KView on X ↗
strongly agree that we need "a much broader and more ingenious stable of evaluations." getting there will require eval pluralism: more benchmarks, more evaluators (including subject-matter experts), and open eval infrastructure - more benchmarks: a single benchmark measures a slice of what these systems can do, offering an empirical tool to understand capabilities, risks, and failure modes. earlier this year we shared our view of the axes that need coverage: environment complexity (prompt/response → full worlds), output complexity (fixed-schema labels → nuanced, subjective artifacts), and autonomy horizon (single turns → continually improving agents): x.com/vincentsunnchen/status/202…. @ajratner shared his mental model on 'smoothing' eval coverage: x.com/ajratner/status/2097019446… - more evaluators: critically, measurement shouldn't sit only with a small group of AI researchers. we are encouraged by the increasing number of independent researchers working on evaluation / data, and believe even more subject-matter expertise will be needed to close the evaluation gap. people with the most to gain (and lose!) from these systems should be the ones defining what "good" looks like: clinicians, engineers, scientists, and everyday users - open eval infrastructure: the above requires a shared infrastructure built with people and technology. we love @ryan_marten's framing of "benchmarks are software" (x.com/ryan_marten/status/2080321…) and see room for continued investment in continuous QC, grading design, and reward-hacking monitoring, esp as models show increasing 'situational awareness' (h/t @henryehrenberg: senior-swe-bench.snorkel.ai/blog…) Open Benchmarks Grants (benchmarks.snorkel.ai) is just the start of our push to support all of the above - and we have a lot more coming soon. please reach out if you are building here - we would love to work with you!

Interviews & talks

Connections

Who they amplify

Accounts whose posts Vincent has reposted recently: Chuck Ng, Samuel Pearton, Ishan Mukherjee.

Why it matters here

The most reachable Snorkel person for someone who wants to create a public benchmark: a taste/design-judgement benchmark proposal to the Open Benchmarks Grants is a concrete, credibility-building path for a new specialist.

How to reach

Public on X (@vincentsunnchen) and via the grants program; lead with a crisp benchmark thesis (what capability, why unsaturated, how experts validate). Speculative but aligned with his stated criteria.

What we could not establish
  • Exact degree(s) at Stanford
  • Whether his formal title is co-founder (Fortune) or founding team member (own site)
  • Location

Sources

  1. Snorkel AI – Vincent Sunn Chen author page
  2. Vincent Sunn Chen – personal site
  3. Benchmarks should shape the frontier, not just measure it (Apr 7, 2026)
  4. Snorkel press: Fortune on OpenAI Astra evaluation metrics (Sep 4, 2026)
  5. Snorkel AI highlights first wave of Open Benchmarks Grants projects (Jul 24, 2026)