Prolific's VP of Data & AI (earlier listed as VP of Data, Research & Analytics). He leads the data science behind participant quality, AI-agent detection and evaluation products such as HUMAINE (attribution to HUMAINE is inferred, not confirmed).
medium confidenceUpdated 2026-09-19
Background
Economics and computer science background, with 10+ years on distributed decision systems, ranking and recommendation. He previously worked at Meta (Instagram).
What they run now
Data and AI at Prolific: quality measurement, fraud and agent detection, and human-plus-synthetic evaluation methods.
Career
presentVP of Data, Research & Analytics, then VP of Data & AI, Prolific
—Ranking / recommendation (per MLST), Meta (Instagram)
On the record
Training data will not be fully synthetic · 2025-10
He argues human data stays necessary. Synthetic data should be routed to low-stakes tasks and humans to high-stakes judgements. source
Recent posts
3 posts archived · most engaged first, then the latest
Today we're putting our AI research work out in the open: labs.prolific.com
We started our AI research team less than two years ago with a founding mission to work on the science of evaluation: how to measure AI systems well, grounded in real human judgement at population scale. We have since expanded into alignment and interpretability.
Some of what we've built so far:
- HUMAINE, our demographically stratified preference leaderboard; 51 models, 57,000+ human evaluations across 20 demographic groups
- Alignment leaderboard stress-testing models across 904 multi-turn scenarios on honesty, safety, manipulation, corrigibility, scheming
- 6 papers so far, including a spotlight at ICML 2026, work at ICLR 2026, and more under review
with Nora Petrova, Jerome W., John Burden
P.S. We're hiring, and open to collaborators. Please reach out.
Prolific is at ICLR next week. If you're in Rio, let's meet up.
We're bringing two pieces of research about what changes when you stress-test evaluation methodology:
HUMAINE: a demographically stratified leaderboard for LLMs, built on 23,000+ real human conversations across 22 demographic strata in the US and UK. We show how aggregate leaderboard scores conceal meaningful disagreement. lnkd.in/ecMCtzCJ
The Missing Red Line (ICBINC Workshop): 8 frontier models tested in scenarios where commercial objectives conflict with user safety. They fabricate safety information, dismiss medical risks, and don't become more cautious as harm escalates. lnkd.in/eYwCa5wF
I'm also giving an expo talk on why your models are outgrowing your evaluations.
DM me if you want to meet up - whether that's about evals, human data infrastructure, or joining us. We're also co-hosting an afternoon tea social with Encord, details soon.
#ICLR26#AIResearch
Turns out 'podcast in a helicopter' belongs on the bucket list. No laptops, no distractions, just a conversation with Shachar Meir about data & AI, flying over the English country side. Episode's out now
Accounts whose posts Enzo has reposted recently: Billur Engin.
Why it matters here
He is the person who would know how Prolific measures rater agreement and quality on subjective tasks, which is the core technical problem of taste data. Moderate relevance.
How to reach
Via LinkedIn. A methods-level exchange (inter-rater reliability on aesthetic judgements) is the likeliest hook (speculation).