Co-founder and CEO of Prolific. He took an Oxford academic-participant marketplace and turned it into a human-data platform for AI evals, RLHF and red-teaming. He is a public advocate that human data remains 'where the alpha is', and the face of HUMAINE, Prolific's human-preference leaderboard.
high confidenceLondon / Oxford, UKAt Prolific since 2014Updated 2026-09-19
Background
Physicist and computational biologist: BSc physics (University College Cork), MPhil computational biology (Cambridge), and a DPhil in genomic medicine & statistics at Oxford (Wellcome Trust). There he built Mykrobe predictor, an antimicrobial-resistance tool published in Nature Communications (2015) and adopted by Public Health England.
He co-founded Prolific during his PhD in 2014 via Oxford University Innovation's Startup Incubator (the IP was not part of the PhDs). It grew from a two-founder startup serving behavioural researchers into a platform that AI labs use.
What they run now
Moving Prolific up from academic surveys to AI data: evals, RLHF, red-teaming, a programmable 'humans behind an API', and HUMAINE as a vendor-neutral human-preference benchmark. He is also defending Prolific's data-integrity positioning (ID verification, authenticity checks) as AI agents threaten crowdsourced research.
Career
2014 – presentCo-founder & CEO, Prolific
—DPhil researcher, University of Oxford (Wellcome Trust Centre for Human Genetics)
—University College Cork, BSc Physics
—University of Cambridge (Fitzwilliam College), MPhil Computational Biology
—University of Oxford, DPhil Genomic Medicine & Statistics
On the record
Human data still matters · 2025-12
He is bullish that human intelligence must stay in the loop: "human data is where the alpha is." source
Evals beyond vibes · 2025-12
He favours orchestrating LLM judges with human data and argues that static leaderboards and vibes are no longer enough to evaluate models. source
What users value · 2025-12
He says winning models show consistency and a personality and style that appeal across many user types. source
AI agents in online research · 2026-05
Contamination fears are overstated outside MTurk (≤1% detection); attentiveness is the bigger quality issue. He says "The threat appears to be more hypothetical than empirical." source
Pricing per quality · 2026-05
He argues that platforms should be compared on price per quality data point, not per response. source
Recent posts
46 posts archived · most engaged first, then the latest
Every few weeks we hear people say AI will replace survey respondents. Our research team recently fielded a 21-item survey (N = 996) to test that claim.
What we did: Simulated a real, politically representative US sample using LLMs at different levels of fidelity: from no context, to demographic personas, to full interview transcripts.
What we found: Adding demographic-only personas when generating synthetic data roughly tripled distributional error vs. simply asking a model to guess a response with no demographic context. Providing idiographic information (real interview transcripts) alongside demographic personas improved performance slightly, but the biggest improvement came from asking the model to output a distribution of responses rather than picking a single answer (recovering 35-48% of that lost accuracy).
Synthetic data can have a role in research; but this shows it remains a predictor of opinion rather than a replacement for real human insight.
Nice work Andrew Gordon, Nora Petrova, John Burden, Oriol Bosch Jover, and Ning Ding.
Excited to announce that Prolific has been named in Tech Nation's Future Fifty cohort for 2026.
Future Fifty is Tech Nation's recognition of the UK's strongest scaling companies, having supported over 35% of the country's unicorns to date, so I'm pleased to see Prolific part of this year's group alongside companies like Encord pushing AI infrastructure forward.
I'm at the Forum today alongside tech founders, investors, and policymakers. Looking forward to good conversations with all of them.
We're also featured in the Tech Nation 'UK Tech Companies to Watch' report - link in the comments to read.
Very important paper from Sean Westwood demonstrating that agents can now effectively mimic human participants in online data collection.
I'm more optimistic, though. These challenges are tractable. Ensuring we have high integrity, authentic, auditable data collection will force all parties to raise the bar.
We've been proactively adding to our suite of authenticity tools - more coming every week - including many of Sean's recommendations:
prolific.com/resources/prolific-…
If you want to work on these problems, or collaborate on research in this area, get in touch. Much more to come in this space!
Today we launched The Signal, @Prolific's new podcast on research evidence, not research tools or trends.
David and Andrew discussed what research quality looks like once you remove the human from the process. Give it a watch!
Today we launched The Signal, Prolific's new podcast on research evidence, not research tools or trends.
Most research content right now is about the tools: which platform, which model, which workflow. We wanted something new, a series that goes straight to the findings and lets researchers argue about what the evidence actually shows.
The first episode features David Rothschild, Economist at Microsoft Research, and Co-chair of American Association for Public Opinion Research (AAPOR)'s task force on AI in survey research, in conversation with our Head of Research Sciences Andrew Gordon.
The core issue applies well beyond surveys; it’s about what research quality looks like once you remove the human from the process.
Give it a watch! lnkd.in/eTtH2SGa
Our team recently tested whether AI can replace human survey respondents, using simulated LLM personas against a real US survey (N = 996).
Turns out demographic personas make it worse, idiographic information helps a little, and output format is the biggest lever of all.
Synthetic data predicts human opinion; it's not a substitute for measuring it.
Pre-print: papers.ssrn.com/sol3/papers.cfm?…
Partech / Oxford Science Enterprises — Series A co-leads
Who they amplify
Accounts whose posts Phelim has reposted recently: Tech Nation.
Why it matters here
High relevance. HUMAINE measures exactly what taste data is about (style, personality, trust and preference across demographics), so Prolific is the closest 'generalist' to preference and taste evaluation. It is also European and London-based, reachable from Lübeck. Prolific could be a competitor, a supply channel (representative lay raters to set against Julian's expert raters), or a partner.
How to reach
On X as @Phelimb. He appears at European VC and AI events (Partech salons, Prolific summits). A data-backed angle on where lay preference diverges from expert judgement in visual/design evaluation directly extends HUMAINE and is a plausible hook (speculation). Ask whether Prolific sees demand for expert-panel creative evals it cannot serve with general participants.
What we could not establish
Prolific revenue (~$350M) is a Sacra estimate only.
Degree years not found.
Whether Prolific has raised since the 2023 Series A.