Background
Researcher at Epoch AI before co-founding Mechanize. Appeared with Tamay Besiroglu on the Dwarkesh Podcast (April 2025), forecasting drop-in remote-worker automation around 2045. In June 2026 he hosted Mechanize's podcast on why evals are hard to build and how bad training data teaches models to write terrible code.
What they run now
Evaluation and task quality, based on his podcast topics; formal title not published.
Career
- 2025 – presentCo-founder, Mechanize
- 2025Researcher, Epoch AI
On the record
Co-wrote that cheap contractor-labelled data is over: progress now needs interactive environments built by full-time domain specialists whose tacit knowledge is the bottleneck. source
Co-wrote that labs should spend more per RL task (then ~$500, expected to rise to several thousand) because cheap tasks waste expensive compute; data and compute are complements. source
Co-signed the founding statement: build environments and evals to enable full automation of the economy, sizing the prize at ~$18T/yr US wages and ~$60T globally. source
Forecast drop-in remote-worker replacement around 2045; argued progress needs compute, infrastructure, data and complementary innovation together, not just cognitive effort. source
Discussed how low-quality training data teaches models bad coding habits and why good evals are hard to build. source
