Background
Trained as a chemist (University of Washington, 2016) and spent his early career in computational biology: genomics work at BGI, then roughly five years at BioAge Labs (from 2017) building ML pipelines over genomics and proteomics data, with published biomarker research on human mortality. He then moved through fintech and startups (Ambient Finance 2022–23, Ivy Natal 2024–25) before joining OpenAI.
At OpenAI he worked on life-science evaluations — he is a co-author of GeneBench-Pro (released 30 June 2026), GeneBench and LifeSciBench, and contributed to Humanity's Last Exam. That benchmark work is what his company is built on: he saw frontier models reach only about 30% success on bioinformatics tasks.
What they run now
Building a company that produces two kinds of RL data for frontier labs: long-horizon scientific reasoning tasks with controlled ground truth (bioinformatics, statistical analysis), and routine lab work including multimodal judgement such as evaluating photos of cell cultures or Western blots. Planned expansion into chemistry, materials science, healthcare and general knowledge work. No name, funding or hires announced as of September 2026.
Career
- 2026 – presentFounder, New RL-dataset company (unnamed) — Biology and statistics datasets for frontier labs
- 2025 – 2026Researcher, OpenAI — Life-science benchmarks: GeneBench-Pro, GeneBench, LifeSciBench
- 2024 – 2025Ivy Natal
- 2022 – 2023Ambient Finance
- 2017 – 2022ML / computational biology, BioAge Labs — ML pipelines for genomics and proteomics
- —Genomics, BGI — Per RuntimeWire
- 2016University of Washington, Chemistry
On the record
LLMs generalise poorly and have 'spiky' capabilities even in heavily funded areas like coding; most economically valuable skills are missing from existing datasets. source
Forecasts that frontier labs will spend more than $100B on targeted data acquisition (no timeframe or method given). source
Models excel at verifiable tasks like maths and code but struggle where 'research taste' and context matter; most work is not easily encoded into a gradable environment. source
Thinks frontier labs are overvalued and stuck in a 'Red Queen's race' of spending; advised colleagues to take liquidity in tender offers. Says he holds ~$700K of OpenAI stock he cannot sell before an IPO. source
Sceptical that recursive self-improvement is a near-term breakthrough. source
