strongly agree that we need "a much broader and more ingenious stable of evaluations." getting there will require eval pluralism: more benchmarks, more evaluators (including subject-matter experts), and open eval infrastructure
- more benchmarks: a single benchmark measures a slice of what these systems can do, offering an empirical tool to understand capabilities, risks, and failure modes. earlier this year we shared our view of the axes that need coverage: environment complexity (prompt/response → full worlds), output complexity (fixed-schema labels → nuanced, subjective artifacts), and autonomy horizon (single turns → continually improving agents):
x.com/vincentsunnchen/status/202….
@ajratner shared his mental model on 'smoothing' eval coverage:
x.com/ajratner/status/2097019446…
- more evaluators: critically, measurement shouldn't sit only with a small group of AI researchers. we are encouraged by the increasing number of independent researchers working on evaluation / data, and believe even more subject-matter expertise will be needed to close the evaluation gap. people with the most to gain (and lose!) from these systems should be the ones defining what "good" looks like: clinicians, engineers, scientists, and everyday users
- open eval infrastructure: the above requires a shared infrastructure built with people and technology. we love
@ryan_marten's framing of "benchmarks are software" (
x.com/ryan_marten/status/2080321…) and see room for continued investment in continuous QC, grading design, and reward-hacking monitoring, esp as models show increasing 'situational awareness' (h/t
@henryehrenberg:
senior-swe-bench.snorkel.ai/blog…)
Open Benchmarks Grants (
benchmarks.snorkel.ai) is just the start of our push to support all of the above - and we have a lot more coming soon. please reach out if you are building here - we would love to work with you!