Co-founder and CEO of Datacurve, a YC W24 coding-data specialist that treats expert supply as a consumer product (Shipd) rather than a labelling operation. One of the youngest founders in the market; Forbes 30 Under 30 per Waterloo.
high confidenceSan FranciscoAt Datacurve since 2024Updated 2026-09-19
Background
Chinese-Canadian, raised in Canada; built a rock-climbing training app in high school and led a student team building a web app supported by Bank of Montreal (per 36Kr/YC). During a co-op at Cohere working on LLM reasoning she saw how scarce high-quality code training data was, which led to Datacurve. She left the University of Waterloo CS program in 2024 (per 36Kr).
What they run now
Runs the company: lab sales, the Shipd engineer community (bounty design, retention) and research outputs like DeepSWE used to market data quality. Has signalled expansion beyond code into finance, marketing and medicine (TechCrunch, Oct 2025).
Career
2024 – presentCo-founder & CEO, Datacurve
—Co-op / intern, LLM reasoning, Cohere
—University of Waterloo, Computer Science (incomplete) — Dropped out in 2024 per 36Kr
On the record
Expert supply as consumer product · 2025-10
Says Datacurve treats contributor experience as a consumer product, not a labelling operation; pay alone does not retain top engineers. source
Data bottleneck · 2025
Argues the bottleneck of large models is the lack of carefully selected, high-quality annotated data. source
Benchmarks hide divergence · 2026-05
Public leaderboards make top models look close; DeepSWE shows where they actually diverge. source
Recent posts
29 posts archived · most engaged first, then the latest
Today we’re releasing DeepSWE, a new standard for agentic coding benchmarks.
On public leaderboards, top models often look relatively close in capability. DeepSWE shows where they actually diverge, reflecting the realistic experience of developers in their day-to-day work.
Today we’re announcing we’ve raised $17.5 million in funding across a $15M Series A led by Chemistry and a $2.7M Seed to accelerate foundation model progress through providing frontier training data for LLMs.
When we first started Datacurve, it came from a simple realization: foundation model progress is limited not just by compute, but by data quality and complexity.
The right data unlocks new capabilities, especially in coding, where accuracy and reasoning matter most.
We’re now proud to partner with the world’s leading foundation-model labs, providing them with high-quality, complex training data that helps push the boundaries of what AI can do.
This is still just the start. Come build the future of technology with us in San Francisco: datacurve.ai/careers
Huge thanks to our incredible team and investors who’ve believed in us since day one and beyond: @garrytan at @ycombinator, @1vnzh from @cohere , @Mark_Goldberg_ from @chemistry_fund, @TheDerrickLi from @AforeVC, @forwarddeploy, @SoheilK, and @shyamalanadkat.
Opus 4.8 results are now on DeepSWE. Thanks for your patience. We wanted to it to be thorough and here are our findings.
Deep dive on the results will drop in the coming days.
Mark Goldberg — Series A lead investor (Chemistry)
Balaji Srinivasan — seed investor
Who they amplify
Accounts whose posts Serena has reposted recently: Mark Goldberg, Y Combinator.
Why it matters here
Her gamified, community-first supply model (bounties, challenges, consumer UX) is the most transferable playbook for recruiting scarce creative experts, who also respond poorly to pure hourly gig framing. Datacurve is code-only today, so it is a model to study rather than a direct competitor, unless it expands into design/front-end judgement.
How to reach
Speculative: active on X; would likely engage on contributor-experience design and benchmark methodology. Ask how Shipd bounty pricing and quality gates work, and whether Datacurve sees demand for UI/front-end taste data adjacent to code.