Ex-Scale AI general manager who joined Handshake AI in June 2025 to own strategy, positioning and the innovation roadmap. He led the Cleanlab acquisition narrative and is Handshake AI's clearest strategic voice on expert quality.
high confidenceSan FranciscoAt Handshake AI since 2025Updated 2026-09-19
Background
At Scale AI he was General Manager and Director of Product & AI for the International/Global Public Sector division, with earlier roles in quality infrastructure and strategic product management. Before that: Bain & Company consultant, international expansion lead at a consumer products company (The Org lists BARK), and a startup founder. BA in Pure Mathematics and MA in Economics from Boston University.
What they run now
Long-term strategy, market positioning and innovation roadmap; data quality via the Cleanlab team.
Career
2025-06 – presentChief Strategy & Innovation Officer, Handshake AI
2025GM & Director of Product and AI, International Public Sector, Scale AI
—Consultant, Bain & Company
—Boston University, BA Pure Mathematics; MA Economics
On the record
Supply advantage · 2025
Argues Handshake's verified student/PhD network lets it find specific experts in minutes where rivals take weeks. source
Experts undervalued · 2025
Says no company in the industry has fully internalised how critical experts are, and that post-training needs strong writing, attention to detail and real domain expertise. source
Data quality · 2026-01
Framed the Cleanlab acqui-hire as buying years of focus on data-quality assessment. source
Recent posts
37 posts archived · most engaged first, then the latest
AI data factories like Handshake, Scale, Surge, and Mercor have a big role to play in the advancement of AI at the frontier and in the enterprise. This is my full thesis. x.com/i/article/2056792438342385…
There is an astounding amount of human <> agent collaboration required to produce great RL environments and tasks.
The skill of authoring this kind of data will increasingly become a critical full-time, high-skilled job, both at data companies and within enterprises.
It's become increasingly clear that most people dramatically underestimate how essential evals for every enterprise workflow within every individual enterprise are going to be
Knowing whether agents actually work on economically valuable tasks in environments as messy as production is a core bottleneck for AI.
Launching ATLAS-Finance, the most realistic finance eval to date. The best model succeeds on 12% of tasks. Proud of the team!
Excited to announce ATLAS - Visual Life Sciences (VIALS), one of our first benchmarks within the ATLAS series. VIALS consists of 161 tasks, and top models perform < 30%.
ATLAS stands for Agentic Taskforce and Labor Assessment Standard, and it's purpose-built to measure AI systems on economically valuable work across every domain and profession.
We noticed that AI models were largely incapable of visual interpretation across key, everyday life sciences workflows - without this capability we can't hope to meaningfully assist in scientific research and drug discovery.
Check it out!
Paper: arxiv.org/abs/2608.21357
Dataset: huggingface.co/datasets/handshak…
Curtis Northcutt — Cleanlab CEO, acqui-hired Jan 2026
Why it matters here
The person who decides which new data verticals Handshake AI enters. A taste/creative vertical would be a strategy-and-innovation call. High relevance.
How to reach
Active on X. Frame a pitch around a quality problem (measuring taste reliability) rather than supply, since he cares about quality differentiation (speculation).