Owns the human-data budget and vendor list
The index · 524 people · 68 organisations
Who's who
Everyone on both sides of the market for expert human judgement: the labs and model builders who buy it, the companies that sell it, and the researchers who decide what “good” looks like.
Looking for the seventy creative-lane contacts and the network graph? Who to talk to →
The companies that sell expert data to the labs, and the 93 people who run them — each with a full dossier.
New entrants & independents
People worth tracking who are not at one of the indexed companies — new founders, stealth companies, operators between roles.
Mercor
Marketplace that recruits vetted domain experts to create training and evaluation data (and RL environments) for frontier AI labs.
Mercor co-founder who built its technology (AI interviewer, matching, platform) and, per Forbes, was elevated to co-CEO in 2026 to lead new enterprise offerings. Billionaire on paper (~$2.2B, Forbes July 2026).
23-year-old co-founder and CEO of Mercor, the largest expert-data marketplace by gross volume ($2B run rate in June 2026). He is the category's loudest public voice, framing expert AI training as 'the largest new job category in history' and pushing rubrics and RL environments as the product labs buy.
Mercor co-founder who ran operations as COO through the pivot into expert data, then stepped back to board chairman in October 2025. Still a ~22% owner and board chair, but no longer runs day-to-day operations.
Former Uber Chief Product Officer hired as Mercor's President in May 2025 — the senior operator brought in to scale a company run by 22-year-olds. Forbes says he oversees hiring and client reporting.
Mercor's CFO, announced in June 2026, as the company prepared a ~$20B raise. An eight-year Tesla finance veteran — a signal Mercor is building IPO-grade finance.
Human-data operator who led human data operations at OpenAI, then became a Managing Director at Mercor. His LinkedIn now shows Google DeepMind, i.e. he appears to have moved back to the buyer side.
Surge AI
Bootstrapped, high-quality RLHF, evaluation and RL-environment data from vetted experts for roughly a dozen frontier labs.
Founded Surge AI in 2020 with a few million dollars of his own savings and bootstrapped it past $1B of revenue with around 100-250 staff. He reportedly still owns about 75%. He is the most prominent voice arguing that human taste and judgement, not volume, decide which models win.
Surge AI's first engineer (Aug 2020), now described as CTO. He built the platform behind Surge's quality-controlled human labelling.
Handshake AI
AI-data arm of the Handshake campus career network, selling PhD and professional expert data and evaluations to frontier labs.
Co-founder and CEO of Handshake, who 're-founded' the 12-year-old career network around AI training data. It went from zero to $1B+ gross in about 16 months and he was named to the TIME100 AI 2026. He is the main public voice framing AI tutoring as a new part-time profession.
Handshake co-founder, listed on its leadership page and board. His current operating role is not stated publicly.
Handshake co-founder, still listed on the leadership page. His current role and involvement in Handshake AI are not public.
Garrett Lord's long-time number two, who joined Handshake in 2015 and has held product, finance (CFO until May 2022), legal and marketing roles. Now President & COO across the combined career-network and AI business.
Ex-Scale AI general manager who joined Handshake AI in June 2025 to own strategy, positioning and the innovation roadmap. He led the Cleanlab acquisition narrative and is Handshake AI's clearest strategic voice on expert quality.
Listed on Handshake's leadership page as Chief Commercial Officer of Handshake AI, i.e. the lab-facing sales lead. The Org lists prior roles at Palantir and Amgen.
Handshake AI's Chief Business Officer, running business development and operations. He spent four years as Head of Product Deployment and Operations at Scale AI — delivery at the largest incumbent.
Listed as Head of AI Research at Handshake AI. He leads a research team of about 14 (directors, fellows, data scientists), which presumably now includes the Cleanlab team acqui-hired in January 2026.
micro1
Vetted expert network (plus datasets, RL evals and robotics data) supplying training/evaluation data to AI labs and enterprises.
25-year-old solo founder-CEO of micro1, the fastest-growing small-cap in expert data ($500M gross run rate, Aug 2026). He is the sector's most outspoken 'America-first' voice (no sales to Chinese labs) and is pushing beyond hours into resellable datasets and robotics data.
micro1's CFO, also listed as VP Government — relevant as micro1 raises a new round and pitches itself as a China-free American data supplier.
Listed as micro1's CRO since October 2022 — the commercial lead through its growth from ~$7M to a $500M gross run rate. Previously at Turing, a direct competitor.
micro1's senior AI/research leader: a Stanford PhD (Andrew Ng, Dan Jurafsky) and serial founder who gives the company research credibility with labs. Inc. describes him as a former Stanford CS professor and ex-Apple.
Turing
Former remote-developer staffing marketplace now selling coding/STEM expert data, RL environments and evals to frontier labs and enterprises.
Founder-CEO of Turing. He took a remote-developer staffing marketplace and repositioned it as a 'research accelerator' for frontier labs. He is the loudest voice arguing that commodity data labelling is dead and that expert, workflow-grounded data and RL environments are what labs now buy.
Turing's technical co-founder. He is an ML researcher who built the vetting and matching system that turned Turing's developer pool into a data supply chain. With Ece Kamar named CTO in Sep 2026, his current title is unclear.
Turing's new President and Global CRO (Jul 2026). He runs all go-to-market with labs and enterprises. He is a veteran of the first generation of this market: he ran Figure Eight (ex-CrowdFlower), the crowd-labelling platform that Appen acquired in 2019.
Turing's CTO since 1 Sep 2026. She spent 16 years at Microsoft Research and ended as Corporate VP leading the AI Frontiers Lab (AutoGen, Magentic-One agent work). She is the most research-credentialed hire in the expert-data vendors this year.
Invisible Technologies
AI training data (RL Data Lab, expert marketplace) for foundation-model builders plus enterprise AI workflow/agent platform Meridial.
Founded Invisible in 2015 as a human-plus-software 'digital assistant' operation and took it to unicorn status. He stepped down as CEO in a restructuring and is now executive chairman.
Former McKinsey senior partner who ran QuantumBlack Labs (about 1,000 engineers). He became Invisible's CEO in January 2025, has since raised $100M+ at a $2B+ valuation, bought WeCP and repositioned Invisible as an enterprise AI platform on top of its training-data business.
Invisible's Chief Go-To-Market Officer. His leadership-page bio says he supports enterprises training the world's largest LLMs.
Invisible's COO per its leadership page, overseeing operations of a business built on 3,000+ operators in 35+ countries.
Co-leads Invisible's RL Data Lab, which builds the data and evaluation infrastructure model builders use to train and benchmark frontier models.
Co-leads Invisible's RL Data Lab, designing RL training and eval data built around real deployment challenges.
AfterQuery
Expert-written post-training data, long-horizon tasks and RL environments for frontier labs, validated by in-house training runs.
23-year-old co-founder and CEO of AfterQuery, which went from YC W25 to a reported $3.2B valuation in about 18 months. He is the public face of the company and has positioned it as a research-led data vendor that proves data quality to labs with its own training runs and published benchmarks.
22-year-old co-founder and CTO of AfterQuery. Owns the software that validates and filters expert data and the internal training pipeline the company uses to prove quality to labs.
Third, lower-profile co-founder of AfterQuery, named by Mateega as one of the three who left post-grad jobs to start the company. Press coverage (Forbes, TechCrunch) mentions only two founders, so his current role is unclear.
Snorkel AI
Stanford-born 'frontier AI data lab' selling expert-built datasets, rubrics, benchmarks and RL environments to AI labs and enterprises.
Co-founder and CEO of Snorkel AI, which he spun out of the Stanford AI Lab's Snorkel weak-supervision project. He has steered the company from enterprise labeling software to expert data, rubrics and evaluation environments sold to frontier labs, and is one of the most articulate public voices arguing the frontier is now 'expert agentic data'.
Stanford CS professor and MacArthur Fellow whose Hazy Research lab produced Snorkel; academic co-founder of Snorkel AI and serial founder (SambaNova, Together AI, Lattice and Inductiv, both acquired by Apple). Influential as the intellectual root of the 'data-centric AI' school that many data vendors now sell.
Snorkel co-founder focused on technical strategy and engineering (The Org lists him as Head of Engineering). Built the original open-source Snorkel library at Stanford.
Snorkel AI co-founder who leads research; her 2025-26 output is squarely on how expert data and rubrics are evaluated (rubric failure modes, automated benchmark design, RLVR in low-data regimes). She is the research owner behind Snorkel's claim that rubric quality is make-or-break.
Founding team member (described by Fortune as a co-founder) who leads Snorkel's evaluation and expert-feedback systems and directs the $3M Open Benchmarks Grants. He is Snorkel's public face on benchmark integrity, quoted in Fortune (Sep 2026) on OpenAI changing eval metrics around a launch.
Chief Scientist at Snorkel AI and assistant professor of CS at UW-Madison; drives Snorkel's benchmark and evaluation research (OSWorld 2.0, Senior SWE-bench, SlopCode Bench, RIFT) and sits on the Open Benchmarks Grants steering committee.
Labelbox
Sells RL environments, expert data (via its Alignerr network) and an enterprise agent platform to frontier AI labs and enterprises.
Co-founder and CEO of Upcraft (Chicago, AI sales agents) which Labelbox acquired on Feb 10, 2026 to automate recruiting, qualification and engagement of Alignerr experts; now listed as Labelbox VP of Growth. Effectively owns the machinery that grows Labelbox's expert supply.
Co-founder and CEO of Labelbox, which he has moved from labeling software to a 'data factory' for frontier labs built on the Alignerr expert network, and in 2026 to RL environments (Horizon) and an enterprise agent platform (Recursion). He controls one of the largest expert networks in the market and is an active acquirer (Upcraft, Feb 2026).
Labelbox co-founder, variously listed as President (Crunchbase), COO (First Round Review) and Chief Product Officer (Craft). An aerospace engineer who met Manu Sharma at Embry-Riddle.
Toloka
Amsterdam-based AI data company selling expert-built training, evaluation and agent data to labs, with Mindrift as its expert network.
Founded Toloka inside Yandex in 2014 and has run it as CEO since 2020, through its carve-out via Nebius and the 2025 Bezos-led round that gave it independent voting control. She also leads Mindrift, Toloka's expert-sourcing platform, so she directly owns both the lab-facing business and the expert supply side.
Former Yandex CTO and ex-CEO of Microsoft's Advertising and Web Services (Bing), who became Shopify CTO in 2024 and joined Toloka as executive chairman when he co-invested in the May 2025 Bezos-led round. Gives Toloka senior big-tech credibility with lab and enterprise buyers.
CTO of Toloka per The Org (unverified profile), leading engineering for the data platform. Ex-Yandex.
Chief Sales Officer at Toloka per The Org (unverified), leading US sales; previously at Turing, a direct competitor in expert data for labs.
Appen
ASX-listed, 30-year-old crowd data vendor (1M+ contributors) selling annotation, LLM expert data and evals to hyperscalers and model builders.
Appen's CEO since Feb 2024. He took over days after Google's contract termination and has run the turnaround: cost cuts, a pivot to generative-AI and LLM expert data, and strong growth in Appen China. He is the operator of the sector's only audited, public pure-play.
Appen's CFO (interim from Aug 2023, permanent from Feb 2024). He has been at Appen since 2016 and is the keeper of the only audited P&L among the pure-play human-data vendors.
Appen's product and technology chief. He owns CrowdGen (the crowd's work interface), Mercury (the project and crowd backend) and the expert-sourcing platform Elite AI.
Appen Global's sales lead since Dec 2024, recruited from Scale AI and Snorkel. He is the GTM face to US labs and hyperscalers outside China.
Head of Appen's crowd (1M+ contributors, 200+ countries) since Aug 2025, which makes him the supply-side owner. He took over as crowd satisfaction fell (Crowd NPS dropped from 33 to 22 in FY2025).
Runs Appen China, now nearly half of Appen's revenue (US$102.9M in FY2025, +75%) and its profit engine. Appen describes it as the largest AI data company in China.
Deccan AI
India-operated post-training data, RL environments (STARK) and evals (Helix) for frontier labs and enterprises.
Founder of Deccan AI, an India-operated expert-data vendor that raised a $25M Series A in March 2026. A former banker/consultant/strategist who pitches 'accuracy' over superintelligence and runs a very large Indian expert pool at lower cost than US rivals.
Deccan AI's revenue/sales leader, listed on the company leadership page with 20+ years of B2B sales at AWS and Intuit and a Kellogg MBA. Owns the lab and enterprise pipeline.
Heads strategy at Deccan AI per its leadership page; 15+ years across Google, McKinsey and Amex, IIT Madras and IIM Ahmedabad.
Contra Labs (Contra.Work Inc.)
Creative human-data unit of the Contra freelance marketplace: designer preference data, trajectories, benchmarks and a creative arena for AI labs and tools.
Co-founder and CEO of Contra, the commission-free freelance marketplace (about $45M raised; Unusual, NEA, Cowboy). In March 2026 he repointed Contra's creative supply into Contra Labs, the most direct incumbent in creative human data. A creative himself (audio engineer, music composer), he pitches Contra Labs as 'the eval layer for creative AI'.
Co-founder and CTO of Contra; the engineering lead behind the marketplace that supplies Contra Labs. A well-known open-source JavaScript/PostgreSQL engineer (Slonik and other libraries) who blogs on engineering and, since 2026, on AI-assisted development and fine-tuning.
First author of Contra Labs' flagship paper, The Human Creativity Benchmark (Jun 2026), listed with a dual Contra (New York) and MIT affiliation. She is the most credentialed ML researcher publicly attached to Contra Labs, and she may be the unnamed 'Head of Research', which is unconfirmed.
Contra Labs researcher named on both of its public papers: the Human Creativity Benchmark and the TASTE designer-preference dataset with Lica World. She is one of the few people who has run professional-designer annotation studies at commercial scale and published the method.
Taste Labs
Sells design preference data, rubrics and RL/eval environments to frontier labs, plus brand-verification software (Brand API) to app companies.
Founder and CEO of Taste Labs, the best-funded pure-play 'taste' data company ($18.5M seed from CRV and Amplify, Jun 2026). She came from growth, strategy and marketing at Exa, a search-API startup selling to AI labs, so she has sold to labs before. She is the category's main public voice, arguing that AI slop is caused by post-training on averaged preferences.
Research-side Member of Technical Staff at Taste Labs. She wrote the company's public research agenda and runs its Prototype fellowship for outside researchers. Before Taste Labs she worked on the economics and provenance of AI training data, a rare background in this market.
Gray Swan AI
AI security firm: a 15,000-person prize-paid red-teaming crowd (Arena) feeding automated red-teaming (Shade) and runtime guardrails (Cygnal) sold to labs and enterprises.
CMU security and privacy professor who runs Gray Swan, the company that turned a crowd of prize-paid jailbreakers into a frontier-lab red-teaming and guardrail business. He is the operator of one of the few human-data crowds paid by competition rather than by the hour.
AI-safety researcher behind the universal adversarial-suffix jailbreak, representation engineering, circuit breakers and HarmBench, and co-founder of Gray Swan. His research is the technical core of the company's attack and defence products.
Head of CMU's Machine Learning Department, OpenAI board member chairing its Safety and Security Committee, and Gray Swan's chief scientist. Few people sit closer to how frontier labs think about adversarial testing.
Enterprise-security go-to-market executive hired in January 2026 to lead Gray Swan's global market expansion. He is the commercial counterpart to the academic founders.
Listed by Gray Swan as Chief Product Officer, owning the Shade/Cygnal/Arena product line. Little public detail verified.
Centaur Labs (trading as Centaur AI)
Expert medical/scientific data annotation and model evaluation via a gamified expert crowd (DiagnosUs), sold to health and AI developers.
MIT-trained collective-intelligence researcher who turned a thesis finding (crowds of medical students can out-classify individual dermatologists) into Centaur, the best-known gamified expert-annotation company in medicine. His core idea is performance-weighted expert crowds rather than credential-weighted ones.
Centaur co-founder and VP of Engineering with prior experience running mapping and data-labeling teams at Cruise, i.e. hands-on operational knowledge of large labeling programs.
Self-taught engineer and Centaur co-founder, listed as CTO; built the platform behind DiagnosUs and the expert-aggregation pipeline.
Mechanize
Builds RL environments and evaluations (software-engineering tasks with automated graders) for frontier labs' coding agents.
Co-founder of Epoch AI who left to found Mechanize, a small RL-environment shop reportedly in $1.5B+ licence-and-hire talks with Google. He is the most prominent advocate of the view that frontier training data must be expensive, expert-built and sold as artefacts, not hours.
Former Epoch AI researcher and Mechanize co-founder who speaks publicly on evals and training-data quality. Known for long-timeline, economics-first views of AI progress.
Former Epoch AI researcher and Mechanize co-founder, co-author of its essays on RL-task economics and the automation of labour. A prominent, contrarian voice on AI economics and AI-risk debates.
Sepal AI
Former YC S24 vendor of expert-built benchmarks, evals, training data and RL environments for labs; acquired by Mercor.
Co-founder and CEO of Sepal AI, sold to Mercor in February 2026. Before Sepal he helped scale Turing's LLM-trainer business, so he has run expert supply at two of the market's vendors.
Technical co-founder of Sepal AI; after the Mercor acquisition his site says he is working on RL environments and links to Mercor.
Sepal AI co-founder who ran go-to-market and operations; earlier built Turing's foundational LLM-trainer business and managed operations for 500+ AI trainers. Her LinkedIn headline now shows Mercor.
Datacurve
Expert coding data, RL environments and evals for frontier labs, sourced via the gamified Shipd bounty platform for engineers.
Co-founder and CEO of Datacurve, a YC W24 coding-data specialist that treats expert supply as a consumer product (Shipd) rather than a labelling operation. One of the youngest founders in the market; Forbes 30 Under 30 per Waterloo.
Technical co-founder of Datacurve and co-author of its DeepSWE benchmark. Former Google intern and Waterloo CS student focused on multimodal RL and browser-automation agents.
Halluminate
RL environments and expert-built data training frontier models on investment banking, private equity and consulting work.
Ex-Capital One Labs product/research lead who co-founded Halluminate (YC S25), first a computer-use agent benchmark and simulation company, now selling finance RL environments built by ex-bankers and deal professionals to frontier model developers.
Halluminate co-founder/CTO, a Cornell CS graduate and Milstein scholar with data-engineering experience at two NYC startups; co-author of the company's 2026 finance due-diligence benchmark.
Irregular (formerly Pattern Labs)
Frontier AI security lab running pre-release offensive-cyber evaluations (SOLVE, CyScenarioBench, FrontierCyber) for OpenAI, Anthropic, Google DeepMind and Meta.
CEO of Irregular, the Tel Aviv firm that has become the default third-party cyber evaluator named in frontier-lab system cards. He sells expert-built evaluation scenarios, not labour hours, to the most competitive buyers in the market.
CTO of Irregular, responsible for the simulated attack-and-defence environments that frontier labs use for pre-release cyber evaluations.
Vals AI
Independent, private-test-set benchmarks of AI models on real professional work (law, finance, medicine, coding), sold to labs, vendors and enterprises.
25-year-old Stanford CS graduate who co-founded Vals, now the most-cited independent benchmark shop for professional-domain AI (law, finance, medicine), with results in lab model cards. He is the face of the 'independent scorekeeper' thesis that won a16z's $40M Series A in August 2026.
Stanford CS graduate and Vals co-founder/CTO who builds the evaluation infrastructure behind Vals' private benchmarks; background in quant (Hudson River Trading) and big-tech engineering internships.
Design Arena (Intelligence, formerly Arcada Labs)
Free crowdsourced arena where millions vote on AI design outputs; sells the preference data and custom evals to frontier labs.
Co-founder and CEO of Intelligence, the company behind Design Arena, the largest crowdsourced preference arena for AI-generated design. She claims $60M ARR selling arena data and evaluations to frontier labs, which makes her the commercial proof point, and the main counter-thesis, for Julian's lane: consumer-crowd taste instead of expert panels.
Co-founder and CTO of Intelligence / Design Arena. She owns the arena stack: prompt routing to many models, vote collection and the leaderboards across web, image, video, audio, slides and agent categories. That stack produces the data sold to frontier labs.
Arena (formerly LMArena / Chatbot Arena; Arena Intelligence Inc.)
Crowdsourced model leaderboard whose free votes power paid evaluations and preference data sold to labs and enterprises.
Statistician-turned-CEO who runs Arena, the crowdsourced leaderboard that turned into a ~$100M run-rate evaluation business in under a year. He is the person who sets the terms on which free crowd preference competes with paid expert judgement, the central question for any specialist human-data vendor.
Berkeley professor and serial founder (Databricks, Anyscale) who advised the Chatbot Arena students, co-founded the company and chairs it. He gives Arena its enterprise credibility and company-building playbook.
Systems researcher who started Chatbot Arena as a Berkeley PhD 'weekend project' and now runs Arena's technology. He built the infrastructure (FastChat, Vicuna, MT-Bench, Arena-Hard) that made crowdsourced pairwise preference the default public yardstick for models.
Prolific
Self-serve marketplace of ID-verified research participants and AI taskers, sold to academics and AI teams for evals, RLHF and red-teaming.
Psychological scientist who co-founded Prolific after struggling to recruit participants for her own PhD studies. She was the company's early CEO. Current public materials name only Phelim Bradley as CEO/co-founder, and her present role is unclear.
Co-founder and CEO of Prolific. He took an Oxford academic-participant marketplace and turned it into a human-data platform for AI evals, RLHF and red-teaming. He is a public advocate that human data remains 'where the alpha is', and the face of HUMAINE, Prolific's human-preference leaderboard.
Prolific's VP of Data & AI (earlier listed as VP of Data, Research & Analytics). He leads the data science behind participant quality, AI-agent detection and evaluation products such as HUMAINE (attribution to HUMAINE is inferred, not confirmed).
Prolific's VP Product, owning the product that packages 'verified, diverse humans behind an API' for AI teams. She is one of the few voices in this market arguing publicly that models need cultural sensitivity and human judgement, not just more data.
Needs the data for a model; writes the spec
Runs vendor projects; decides if a pilot converts
Contracts, licensing, provenance
Decides whether the problem is funded
A talk to first · B map and watch · C context. Tier C is hidden by default.
OpenAI
Frontier lab behind ChatGPT, GPT-5.x and the GPT Image family (Images 2.0 Apr 2026, 2.5 Sep 2026 with templates/posters/sketch). The biggest single buyer of expert human data, now with image/design use cases shipping in ChatGPT.
'Product Lead' in ChatGPT Images credits (Dec 2025); Images 2.0/2.5 add templates, posters, sketch editing (design-adjacent)
Listed under 'Data & Evaluation' in ChatGPT Images (gpt-image-1.5) credits (Dec 2025); GDPval co-author (Oct 2025); LinkedIn headline 'OpenAI | Human Data'
Research Lead in ChatGPT Images credits (Dec 2025); image generation leadership for GPT-4o image gen (Mar 2025)
Credited 'Data Leads' for GPT-4o image generation (Mar 2025); GDPval co-author (Oct 2025)
GPT-4o system card 'Data acquisition leads' (2024); listed in Data & Evaluation for ChatGPT Images (Dec 2025)
Credited 'Data Leads' for GPT-4o image generation (Mar 2025) and in Data & Evaluation for ChatGPT Images (Dec 2025)
Human Data lead on GPT-4o system card (2024); co-author of GDPval (expert-graded eval); named VP Research & Safety, safety teams report to her after Heidecke exit (Jul 2026)
'Multimodal Lead' in ChatGPT Images credits (Dec 2025); 'Multimodal Organization' lead for 4o image gen
First author & correspondent on GDPval, expert-graded eval across 44 occupations built with industry experts and vendor partners (Oct 2025)
'World Simulation Lead' in ChatGPT Images credits (Dec 2025); Sora 2 'Organization' lead (Sep 2025). Sora app shut down Mar 2026 - current remit unverified
Listed under 'Data & Evaluation' in ChatGPT Images (gpt-image-1.5) credits
Listed under 'Data & Evaluation' in ChatGPT Images (gpt-image-1.5) credits
Listed under 'Data & Evaluation' in ChatGPT Images (gpt-image-1.5) credits
Listed under 'Data & Evaluation' in ChatGPT Images (gpt-image-1.5) credits
Listed under 'Data & Evaluation' in ChatGPT Images (gpt-image-1.5) credits
Equal-contribution author on GDPval expert-graded eval (Oct 2025)
Listed under 'Data & Evaluation' in ChatGPT Images (gpt-image-1.5) credits
Listed under 'Data & Evaluation' in ChatGPT Images (gpt-image-1.5) credits
Listed under 'Data & Evaluation' in ChatGPT Images (gpt-image-1.5) credits
Listed under 'Data & Evaluation' in ChatGPT Images (gpt-image-1.5) credits
Credited 'Model Behavior' for 4o image generation (Mar 2025) - taste/behaviour spec for image outputs
'Human Data Advisors' on 4o image gen (Mar 2025); InstructGPT/RLHF first author; GPT-4 IF data collection lead
Chief Research Officer, quoted on safety/research restructuring (Jul 2026); in ChatGPT Images leadership credits (Dec 2025)
Co-author on both HealthBench (262 physician graders, May 2025) and GDPval (Oct 2025) - likely expert-data ops; exact role unverified
Listed under 'Data & Evaluation' in ChatGPT Images (gpt-image-1.5) credits
Equal-contribution author on GDPval expert-graded eval (Oct 2025)
Safety Lead for 4o image gen (Mar 2025); listed in Data & Evaluation for ChatGPT Images (Dec 2025)
Listed under 'Data & Evaluation' in ChatGPT Images (gpt-image-1.5) credits
Core research on 4o image gen (Mar 2025); Listed under 'Data & Evaluation' in ChatGPT Images (gpt-image-1.5) credits (Dec 2025)
Listed under 'Data & Evaluation' in ChatGPT Images (gpt-image-1.5) credits
Listed under 'Data & Evaluation' in ChatGPT Images (gpt-image-1.5) credits
Google DeepMind
Builds Gemini plus the Imagen, Nano Banana, Veo, Genie and Lyria generative-media models; runs very large human side-by-side evals for subjective visual quality, which makes it a top buyer of image/video taste judgement.
Veo co-lead (Google AI Release Notes podcast); discussed video eval and cautioned human preferences are unreliable optimisation target
Authored I/O 2025 generative-media post (Veo 3, Imagen 4, Flow, Lyria 2); Imagen 3 advisor
Nano Banana engineering lead on Sequoia 'Training Data' podcast; credits meticulous data curation; Imagen 3 core contributor
Co-author of Gecko (T2I eval with >100K human ratings) and Imagen 3 core contributor
Product lead for Nano Banana; describes eval stack of 'large-scale human evals for subjective quality', trusted testers and artist reviews; Imagen 3 core contributor
Nano Banana research lead guest on a16z podcast
Lead author of Gecko: T2I evaluation, prompts and human ratings (>100K annotations)
Senior author on Gecko human-rating eval paper
Imagen 3 core contributor; co-author PASTA (preference-adaptive T2I with paid rater data, ICML 2025)
Gecko co-author; Imagen 3 contributor
Co-author of Genie 3 world-model announcement
Imagen 3 core contributor (report details 366K ratings from 3,225 raters via external platform)
Took over Gemini app from Sissie Hsiao (Apr 2025); Labs owns Flow/creative tools
Named Nano Banana team member
Named Nano Banana core team member; Imagen 3 core contributor
Named as working on Gemini RL and Omni Thinking on AI Engineer leads panel with Veo/Nano Banana leads
Co-author of Genie 3 announcement
Imagen 3 advisor; believed to lead Gemini model product (not verified this pass)
Gecko co-author; Imagen 3 core contributor
Anthropic
Builds Claude. Runs an in-house Human Data team plus vendor pipelines; launched Claude Design (Apr 2026), which makes UI/visual design taste an explicit training and eval target.
Leads design for Claude (ex-Director of Design, Figma); talks publicly about AI and taste/judgement.
Moved from CPO to co-lead Labs in Jan 2026; Labs ships Claude Design (Apr 2026).
Wrote Anthropic's frontend-design grading rubric (design quality, originality, craft, functionality) for long-running app agents; co-authored the frontend-design Skill post (Nov 2025).
Leads character/personality fine-tuning and wrote Claude's Jan 2026 constitution; defines what 'good' outputs look like.
Named as leading Labs with Krieger, Jan 2026.
Started and leads Societal Impacts (human studies, values, red-teaming datasets).
Co-author of 'Teaching Claude Why' on safety training data (May 2026).
Listed as Chief Science Officer; research leadership over training.
Co-author of the Jan 2026 agent-evals guide (human graders used to calibrate LLM graders).
Co-author of 'Improving frontend design through Skills' on generic AI-looking UI.
Co-author of 'Demystifying evals for AI agents', which covers SME review and crowdsourced human grading (Jan 2026).
Meta (Meta Superintelligence Labs)
MSL builds the Muse model family (Muse Spark LLM, Apr 2026; Muse Image and Muse Video, Jul 2026) for Meta AI, Vibes and ads. It is a major buyer of image/video aesthetic and preference data.
Hired for GPT-4o image generation work; invented text-to-image architectures. Current role not confirmed after 2025.
Led development of Muse Image and Muse Video (Jul 2026); ex-OpenAI Perception lead.
Oversees the technical roadmap for the full Muse series, including Muse Image and Muse Video; described as a synthetic-data leader.
Hired as GPT-4o co-creator with post-training leadership. Evidence is from 2025.
Named in SemiAnalysis Jul 2026 as an ex-OpenAI MSL hire.
Ex-Gemini pretraining/reasoning; listed among MSL researchers in Jul 2026.
Named in SemiAnalysis Jul 2026 as an ex-OpenAI MSL hire (RL/reasoning).
Leads the product arm that ships Meta AI and Vibes features. Role confirmed Aug 2025; not confirmed in 2026.
Hired for Gemini post-training, coding, reasoning and perception. Evidence is from 2025.
Hired as GPT-4o voice co-creator and multimodal post-training lead. Evidence is from 2025.
RL on chain of thought; listed among MSL researchers in Jul 2026.
Named in SemiAnalysis Jul 2026 as an ex-OpenAI MSL hire; prior RLHF/alignment publications.
Ideogram
Text-to-image lab (ex-Google Imagen founders) specialised in typography and graphic design; Ideogram 4.0 (June 2026) open-weights. Explicitly benchmarks against graphic-designer preference, so a natural buyer of designer pairwise judgements.
Co-founder, Imagen lead author; image model research.
Co-founder of Ideogram (2022), ex-Google Brain Imagen; would sign off on designer-preference eval spend.
Founding designer on AI — likely defines design-quality rubrics and internal taste evals.
Co-founder, ex-Google Brain Imagen co-author; research lead on model training.
ML researcher at Ideogram per ZoomInfo directory.
ML team MTS per ZoomInfo directory.
ML MTS per ZoomInfo directory.
Co-founder, diffusion-model pioneer (DDPM); current day-to-day role not verified.
Creative lead — probable internal judge of aesthetic output quality.
Krea
Creative-tool platform turned model lab (FLUX.1 Krea, Krea 2 open weights, Krea Realtime video). Openly 'opinionated' about aesthetics and anti-'AI look'. Already commissioned a Contra Labs creative-expert benchmark. One of the most taste-focused buyers.
CTO and co-founder per Contrary (June 2025).
Second author of the Krea 2 report and co-author of the FLUX.1 Krea aesthetics/preference post.
First author of the Krea 2 report and co-author of the FLUX.1 Krea post-training write-up on opinionated preference data.
CEO; listed on the Krea 2 technical report; sets the aesthetic direction.
Author on Krea 2 technical report (June 2026).
Luma AI
Video/world-model lab (Ray3, Ray3.2, Uni-1.1 image model) backed by HUMAIN (Series C, Nov 2025). Publishes human-rated evaluation reports built with filmmakers and artists, a direct fit for expert creative-judgement data.
CEO; says post-training should 'show it good examples', so quality curation matters to him.
Chief Scientist; publicly discusses data curation and ratios; senior author on TVM (2025-11).
Co-authored the Ray3 human evaluation report (trained expert evaluators, creative-industry framework). LinkedIn lists her at Luma AI.
Co-authored the Ray3 human evaluation report. LinkedIn lists her at Luma AI.
Co-founder/CTO per The Org (listed as 'Alex Y.'; unverified).
Lead author of Terminal Velocity Matching (Luma, 2025-11) and Inductive Moment Matching.
Runs Luma's creator studio in LA, which is a source of expert creative feedback.
Black Forest Labs
Makers of FLUX (image, editing, and since July 2026 FLUX 3 joint image/video/audio). Stable Diffusion's original authors. ~$3.25B valuation, Meta/Adobe/Canva customers. Runs pairwise human-preference evals and is building a data procurement function. A top taste-data target near Julian.
BFL co-founder (Stable Diffusion/SVD author); FLUX.1 Kontext author; senior research leader. Exact current title not verified.
BFL co-founder (latent diffusion/rectified-flow author); FLUX.1 Kontext author. Exact current title not verified.
CEO quoted at the FLUX 3 launch, which was announced with human-preference win rates; final say on data/eval spend.
Co-author of FLUX.1 Kontext paper (June 2025) and known for adversarial diffusion distillation (LADD, which the post-training job also names); likely post-training/distillation lead.
Co-author of FLUX.1 Kontext paper (June 2025); original latent diffusion co-author.
Co-author of FLUX.1 Kontext paper (June 2025); SDXL lead author.
Stanford page lists him as co-founder and research scientist at BFL; Kontext author.
Co-author of FLUX.1 Kontext paper (June 2025); senior research staff.
Canva (incl. Leonardo.ai)
Design platform that built its own layered 'design foundation model' (Oct 2025, Canva AI 2.0) and owns Leonardo.ai (Phoenix, Lucid Origin image models). Buyer for design/layout quality judgements and image aesthetics.
Owns Canva AI; described design-model training on professional templates.
Leonardo CEO post-acquisition; led Phoenix and Lucid Origin foundation models.
Head of AI Research job posting says role reports to him as Gen AI lead.
CTO; Head of AI Research reports to him.
CPO; public face of Canva AI 2.0.
Leonardo co-founder; current role not verified.
Leonardo co-founder; current role not verified.
Midjourney
Leading aesthetics-first image and video model (V8.1 in 2026), small self-funded team. Gets huge in-house taste signal from user image rankings and personalization. Also has a research group studying creative taste and preference data, a direct fit for Julian.
Founder and primary decision-maker for product and research direction.
Midjourney research scientist on 'aesthetic AI models'. Lead author of LiteraryTaste (a paid pairwise taste dataset) and of the diversity-DPO work.
OpenReview lists him as Researcher at Midjourney; senior author on the creative-taste preference papers.
Midjourney-affiliated co-author of LiteraryTaste (Nov 2025), a creative-writing taste preference dataset gathered from paid Upwork annotators.
Midjourney-affiliated co-author of LiteraryTaste (Nov 2025), a creative-writing taste preference dataset gathered from paid Upwork annotators.
Runway
AI video/world-model lab (Gen-4.5, robotics/world models) serving filmmakers and enterprises, ~$5.3B valuation. Film-industry DNA and its own studio make it a strong buyer of expert cinematic/aesthetic judgement.
Promoted to Co-CEO Feb 2026; long-time research/technical head of Runway's models.
Listed as Research Director at Runway AI Summit 2026; senior owner of model research.
Listed as Research Lead at Runway AI Summit 2026. Research focus not verified.
Co-founder with a design background; CIO per TechCrunch May 2026.
CPO per Runway AI Summit 2026 speaker list.
General Counsel; would own data licensing and vendor contracts.
Co-CEO; design/art background; drives creative-industry positioning.
CCO (Feb 2026). Leads Runway's creative org and is a likely internal arbiter of aesthetic quality.
Named CTO in Feb 2026 leadership announcement.
Adobe (Firefly / Generative AI)
Firefly image/video/audio/vector models marketed as 'commercially safe' — trained on licensed Adobe Stock data. Largest licensed-data buyer in creative AI; pays contributors annually and hires teams for dataset acquisition and annotation.
Leads Adobe's generative AI/Firefly model org per Adobe blog author page.
Quoted as CTO at Firefly launch at Adobe MAX 2025.
Founded Firefly Enterprise and Firefly Foundry (custom enterprise models trained on customer IP).
Pika
Consumer AI video lab (Pika 2.5, Pikaformance, social app, 'AI Selves') valued at $470M. It is hiring its own post-training researcher to run video human-preference studies and build reward models, which is an explicit buyer signal.
CTO and research lead (diffusion researcher); the post-training / preference-study hire would likely report to her.
CEO; controls budget at a small startup.
Founding Creative Director; likely internal judge of output quality and style.
Recraft
Design-focused image/vector generation model (Recraft V3, V4 Feb 2026) trained on its own ~20B-param models; markets 'design taste'. Core buyer profile for designer aesthetic judgements.
Founder/CEO; publicly frames V4 as 'design taste meets AI' — owns the taste positioning.
Named Head of AI in Nebius Recraft V4 case study; runs model training.
Head of Design — likely coordinates the designer collaboration used to tune V4.
ElevenLabs
Leading AI audio lab (TTS, dubbing, Scribe STT, Eleven Music; $11B valuation Feb 2026). Active, explicit buyer of voice data and runs data ops with external labelling vendors; audio not visual, but a template for creative-taste data buying.
Co-founder/CTO leading research on core models.
Runs AI safety (~17 people) — consent/verification of voice data.
Leads key partnerships; possible counterpart for data partnerships.
CEO; Series D announcement Feb 2026.
VP Ops managing ~64 people; possible home of data ops (surname not shown).
Stability AI
Stable Diffusion maker, now pivoted to enterprise and professional creative workflows (image, audio, 3D) with music-label and EA partners. Smaller research bench after departures. Relevant for licensed audio/visual data and brand-adherence evals.
Named CTO per Stability AI announcement.
Senior author of ARC post-training (2025) and Stable Audio Open; LinkedIn lists him as Research Scientist at Stability AI.
CEO since June 2024; drove the enterprise/music-label strategy.
Lead author of Stable Audio Open and ARC co-author; LinkedIn shows Stability AI.
xAI (SpaceXAI)
Builds Grok and Grok Imagine (image and video, ranked #1 text-to-video on Artificial Analysis in Jan 2026). SpaceX acquired it in Feb 2026 and it became SpaceX's AI division in May 2026. It runs a large in-house 'AI tutor' workforce and hires expert tutors directly.
Took over the data annotation (tutor) team in Sep 2025 amid 500 layoffs. Named leader of 'Expert Mentors & Grokopedia' in the Feb 2026 lineup
LinkedIn headline seen in search results (undated). xAI had heavy tutor layoffs in Sep 2025, so current status is unverified
Announced 14 Mar 2026 as joining xAI/SpaceX to lead underlying Grok model work in the rebuild
LinkedIn headline seen in search results (undated). xAI had heavy tutor layoffs in Sep 2025, so current status is unverified
LinkedIn headline seen in search results (undated). xAI had heavy tutor layoffs in Sep 2025, so current status is unverified
LinkedIn headline seen in search results (undated). xAI had heavy tutor layoffs in Sep 2025, so current status is unverified
Duke alumni event billed her as 'Human Data Team Lead, xAI'. She joined as an AI tutor and was promoted in Oct 2024
LinkedIn headline seen in search results (undated). xAI had heavy tutor layoffs in Sep 2025, so current status is unverified
Microsoft AI (MAI)
Suleyman's in-house model lab (MAI-Image-2.5, MAI-Voice-2, MAI-Thinking-1) powering Copilot, PowerPoint and Foundry; explicitly courts creatives for image quality, so a plausible taste-data buyer.
Gamma
AI presentation/website/social design generator ($100M ARR, ~50 staff, $2.1B valuation). Its core quality problem is visual layout and design taste; fine-tunes open models and evaluates frontier models.
Heads AI product; would specify design-quality evals.
Runs AI engineering (~18 reports) that owns evals and fine-tuning.
CEO; SaaStr AI 2026 talk on $100M ARR with ~50 people.
Head of Design — likely taste arbiter for generated layouts.
Figma
Design platform shipping Figma Make (prompt-to-app/UI), First Draft and Figma Weave (ex-Weavy image/video canvas, acquired Oct 2025). Runs human 'taste and judgement' evals with contractors — direct fit for UI/design-quality data.
Designed Figma Make's human-centric eval process incl. contractor taste evals.
CTO; owns AI engineering.
CDO — design craft owner; likely sponsor of design-quality eval standards.
CPO; named in Figma AI eval article.
Vercel (v0)
v0 generates UI/React apps via a composite pipeline (frontier LLM + RAG + custom RFT 'autofixer' models). Trains small models with RL; relevant buyer for UI-quality preference data.
Authored v0 model family (2025) and v0 coding-agent (Jan 2026) posts.
Co-author of v0 composite model family post.
Co-author of v0 composite model family post.
Co-author of v0 composite model family post.
Framer
Website builder with AI site generation and Framer Agents (June 2026) working on the design canvas; founders frame 'taste' as the hard part of AI design. Potential buyer of web-design quality evals.
Lovable
Prompt-to-app builder (Stockholm) running on frontier models (Claude Opus, GPT-5.5, etc.). Buys/evaluates model quality for generated UI; a candidate buyer of UI/design-quality evals rather than training data.
Co-founder; manages ~124 per The Org.
Wrote 'We Gave Our Agent a Vent Tool' on agent self-improvement from production friction.
Owns data function; possible owner of eval data.
Head of Product (~10 reports); likely owns output-quality priorities.
Apple Foundation Models (AFM)
Builds on-device/server AFM models plus ADM 3 image model for Image Playground, Genmoji and Photos editing (WWDC 2026). Apple's design culture and UI-generation research make it a natural buyer of design/UI taste data.
Co-author: 'Improving UI Generation Models from Designer Feedback' (CHI 2026) - 21 designers, 1,500 design annotations
Co-author on designer-feedback UI-generation paper (CHI 2026)
Senior author on designer-feedback UI-generation paper (CHI 2026)
Co-author on designer-feedback UI-generation paper
Leads Apple Foundation Models, ML research and AI safety after Giannandrea retirement
Co-author on designer-feedback UI-generation paper
Named head of AFM team (successor to Ruoming Pang)
Mistral AI
Europe's frontier lab (open-weight LLMs, Pixtral vision, Voxtral speech, Le Chat, Forge). Raised €3B Series D (Sep 2026, >€21B valuation, Samsung-led). EU anchor buyer; Le Chat image generation is outsourced to Black Forest Labs FLUX.
Heads the Science org that houses the Human Data Annotation team; author on Magistral, Voxtral, Pixtral
Core contributor on Magistral (RL/post-training) and Voxtral (2025); background in multimodal human-preference evaluation — likely specifier of eval/preference data
LinkedIn headline 'Technical Program Manager, Mistral AI' (UK); Mistral TPM roles are the human-data/vendor program owners — team not confirmed
First-listed core contributor on Magistral RL reasoning paper (Jun 2025)
Core contributor on Magistral (Jun 2025)
First core contributor on Ministral 3 report (Jan 2026) and first author of Voxtral (Jul 2025)
CEO; controls budget after €3B Series D (Sep 2026)
Core contributor on Ministral 3 (Jan 2026) and Magistral (Jun 2025), author on Pixtral
LinkedIn post advertising Mistral's 'Data Annotation Technical Program Manager' role (~May 2024); current status unverified
Core contributor on Magistral (Jun 2025); named in Ministral 3 contributors (Jan 2026)
Core contributor on Magistral (Jun 2025); computer-vision background
Core contributor on Voxtral speech model (Jul 2025)
Core contributor on Ministral 3 report (Jan 2026)
NVIDIA (Nemotron / Cosmos)
Open-weights Nemotron LLMs and Cosmos world/video models; publishes its human preference datasets (HelpSteer) openly, making it a transparent, repeat human-data buyer - mostly text today, video via Cosmos.
Listed under Leadership, Data and Evaluation/Safety in Nemotron 3 report - check if she owns data sourcing
Leads Nemotron post-training teams; senior author on HelpSteer3 (Scale AI + Translated annotators)
Lead author HelpSteer3-Preference: 6,400+ annotators from 77 countries via Scale AI and Translated
Nemotron 3 Leadership list
Listed under Leadership and Data (and as Eileen Peters Long in Evaluation) in Nemotron 3
HelpSteer3 co-author and Nemotron 3 'Data' contributor
HelpSteer3 co-author; Nemotron 3 data and post-training
HelpSteer3 co-author; Nemotron 3 post-training
Nemotron 3 Leadership list
Group builds text2image/text2video/text2-3D foundation models for NVIDIA AI Foundry
HelpSteer3 co-author; Nemotron 3 post-training
Nemotron 3 Leadership and Data lists
Cohere
Enterprise/sovereign LLM company (Command A+, Command A Vision, Aya Vision, North). Acquired Germany's Aleph Alpha (Apr 2026) — a European foothold. Multimodal (document/vision) but enterprise-first; taste data demand mainly UI/document/visual-reasoning.
First author of Aya Vision (May 2025): recaptioning data pipeline, AyaVisionBench and m-WildVision human-preference-style evals
Senior author on Aya Vision (May 2025); Aya multilingual data programmes
Author on Aya Vision and Command A (2025); RLHF research
Joined Aug 2025 as first Chief AI Officer leading research and product development; TIME100 AI 2026
Promoted to head Cohere Labs Sep 2025 after Sara Hooker left; multilingual data/eval research
Command A author (Apr 2025), cited in its human-annotation/adversarial data section (Bartolo et al.)
Chief Scientist per Wikipedia; author on Aya Vision (May 2025)
Second author of Aya Vision (May 2025)
Thinking Machines Lab
Mira Murati's lab (~200 staff): Tinker fine-tuning platform plus Inkling, a natively multimodal (text/image/audio/video) open-weight model (Jul 2026). Pitch is human-in-the-loop customisation — a natural buyer of expert preference data.
RLHF/PPO pioneer; remaining co-founder running research incl. post-training; authored 'LoRA Without Regret' (Sep 2025)
Lead author of 'On-Policy Distillation' post-training blog (Oct 2025)
CEO; public 'human collaboration' thesis ('The Future Worth Building Is Human', Jul 2026)
Named CTO 14 Jan 2026 after Zoph's exit (ex-Meta FAIR/PyTorch)
Amazon AGI (Nova)
Amazon's foundation-model org (Nova family incl. Canvas image and Reel video). Jul 2026 pivot: Canvas/Reel/Omni to keep-the-lights-on, focus on a frontier model from Pieter Abbeel's FMR - image/video taste demand likely reduced.
Leads agent-model training research on AGI team (computer-use/UI agents)
Named by Jassy to lead AI models + AGI org, reporting to CEO
Heads FMR, now Amazon's top AI priority; new frontier model planned for re:Invent 2026
Reflection AI
Ex-DeepMind founders building 'America's open frontier lab' — open-weight frontier LLM/agents (unreleased as of Jun 2026), $25B valuation (Apr 2026), SpaceX compute deal. Text/agentic focus; image/design data demand not evident.
Safe Superintelligence Inc.
Ilya Sutskever's secretive lab (~50 staff mid-2025) pursuing 'straight-shot' superintelligence; no product or model released. Deep-pocketed (Nvidia $5B, Jul 2026) but opaque — low near-term fit for creative/taste data.
Abridge
Clinical ambient-documentation AI (270+ health systems) now training its own models on clinician edits. Expert (clinician) feedback buyer; no visual/design relevance.
No one here matches the filters.
Cognition (Devin)
Devin agent + Devin Desktop (ex-Windsurf); post-trains SWE-series coding models (SWE-2 on Kimi K3, Sep 2026). RL-environment buyer more than human-preference buyer.
Cursor (Anysphere)
AI code editor (>$1B ARR, Nov 2025) training its own Composer models via RL on Kimi K2.5. UI/front-end generation quality is a latent design-taste need, but data comes mostly from real user sessions.
Harvey
Legal AI platform now post-training its own models (Harvey Tenet, Aug 2026). Proven expert-data buyer (Mercor) — but legal text, not creative/design.
Perplexity
AI answer engine/browser (Comet) that post-trains Sonar models on open weights. Answer-quality evals are the data need; limited taste/design relevance.
Sierra
Customer-service agent platform (Bret Taylor, Clay Bavor). Research arm builds τ-bench agent evals; data needs are conversational/agentic, not visual.
ByteDance Seed
ByteDance's foundation-model org: Doubao LLM, Seedream (image), Seedance 2.0 (video, Apr 2026). Among the world's largest image/video-generation efforts — highly relevant demand for aesthetic/video preference data, but access is hard.
DeepSeek
High-Flyer-backed open-weight LLM lab (V3/R1, V4 Flash/Pro 2026). Text/reasoning-first; minimal external human-data buying evident. Context-only for taste data.
No one here matches the filters.
Moonshot AI (Kimi)
Kimi K2/K2.5 (multimodal, MoonViT, Jan 2026) and K3 (2.8T open weights, Jul 2026); ~300 staff, $35B valuation. Its models are used by US firms (Cursor, Cognition, TML bootstrapped Inkling post-training data from K2.5).
No one here matches the filters.
Qwen (Alibaba Tongyi Lab)
Alibaba's LLM/VLM/image-gen family (Qwen3, Qwen3-VL, Qwen-Image 2.0). Qwen-Image RL work (Jun 2026) makes it the most taste-relevant Chinese buyer for image preference data.
Zhipu AI (Z.ai)
Tsinghua spin-out, HK-listed Jan 2026; GLM-5.x LLMs, GLM-4.5V vision, CogView/CogVideo generation. Context-only buyer; multimodal RL work is design-adjacent.
No one here matches the filters.
Computational UI design & models of visual/aesthetic perception of layouts; ERC Advanced Grant 2024–29; CHI Academy 2025
ReproHum project & HEDS 3.0 Human Evaluation Datasheet (2024–25): reproducibility standards for human evaluation
'The Problem of Human Label Variation' (EMNLP 2022) — founding paper on treating annotator disagreement as signal, not noise
Design2Code (NAACL 2025): screenshot-to-webpage benchmark with human evaluation of visual fidelity
LAION-Aesthetics & improved CLIP+MLP aesthetic predictor (2022) — de facto aesthetic filter for Stable Diffusion-era data; audited in 'The Algorithmic Gaze' (2026)
First author GenAI-Arena (NeurIPS 2024) — arena-style human preference collection for generative vision models
PRISM Alignment Dataset (NeurIPS 2024 D&B best paper): participatory, individualised human feedback from 1.5K raters in 75 countries — subjective/pluralistic preference methodology
Senior author on HPSv2 (2023) and HPSv3 (ICCV 2025) human-preference benchmarks for text-to-image
Davidsonian Scene Graph (ICLR 2024, with Google): reliable fine-grained T2I evaluation validated against human judgements
UIClip (UIST 2024) UI design-quality model; 'Improving UI Generation Models from Designer Feedback' (Apple, 2025); DesignPref (Nov 2025)
Senior author UIClip (2024) and DesignPref (2025); long-running crowdsourcing & accessibility research
ImageReward (NeurIPS 2023) and VisionReward (Dec 2024): fine-grained multi-dimensional human preference reward models for image & video generation; also on CogVideoX
Co-lead of RichHF-18K (CVPR 2024 Best Paper) — fine-grained human feedback for T2I; Google team also works on perception/aesthetics modelling
Graphic design generation & evaluation: 'Can GPTs Evaluate Graphic Design Based on Design Principles?' (SIGGRAPH Asia 2024), OpenCOLE (CVPRW 2024), GDUG workshop organiser
CrowdTruth ('Truth is a Lie', 2015) and DICES (NeurIPS 2023 D&B): rater-diversity & 'beyond gold standard' evaluation; also Adversarial Nibbler (T2I safety crowdsourcing)
Toward Verifiable and Reproducible Human Evaluation for Text-to-Image Generation (CVPR 2023): standardised T2I human-eval protocol; showed automatic metrics diverge from humans
HEIM — Holistic Evaluation of Text-to-Image Models (NeurIPS 2023) with crowd human ratings on aesthetics, originality etc.; HELM
TASTE (arXiv 2605.20731, May 2026): designer-annotated 9-dimension preference dataset for AI-generated graphic design, built with Contra (supply side) — closest academic analogue to Miju's thesis
Senior author on RichHF (CVPR 2024 Best Paper); long track record on human visual attention/saliency on UIs and images
Senior author Q-Bench / Q-Align (ICML 2024): LMM visual scoring for quality & aesthetics (IQA/IAA/VQA)
GenAI-Arena (NeurIPS 2024 D&B): open arena collecting human votes on image/video generation & editing; ImagenHub, VideoScore
Human Preference Score v1/v2 (2023) and co-author of HPSv3 (ICCV 2025, with Mizzen AI): wide-spectrum T2I human preference dataset HPDv3
First author DesignPref (Nov 2025): 12K UI design comparisons by professional designers; shows personalised > aggregated preference models
Rich Human Feedback for Text-to-Image Generation (RichHF-18K, CVPR 2024 Best Paper): region-level implausibility/misalignment annotations beyond pairwise prefs
Pick-a-Pic + PickScore (NeurIPS 2023): open dataset of >1M real user preferences on text-to-image outputs
VQAScore & GenAI-Bench (2024): compositional text-to-visual benchmark with large human rating set
First author VBench (CVPR 2024), co-first author VBench-2.0 (2025)
VBench (CVPR 2024) and VBench-2.0 (2025): video-generation benchmark suites with human-preference alignment checks
Writes on art, perception and aesthetics of generative imagery (e.g. 'Can computers create art?', 2018; visual-indeterminacy papers) — theory side of taste
Diffusion-DPO (CVPR 2024): aligning diffusion models directly on Pick-a-Pic human preference pairs — key consumer of T2I preference data
Annotator disagreement / perspectivism & learning from disagreement (2021 survey); MilaNLP hosted Kirk
Senior author Design2Code (2025); broad human-centred NLP and human-eval work
First author Q-Align (ICML 2024) — aesthetic & quality scoring with LMMs; now inside a Chinese frontier lab (overlap with china cluster)
LAION-5B (NeurIPS 2022 D&B) co-author and open-data/scaling-law lead; German open-dataset voice
Co-author of RichHF (CVPR 2024 Best Paper); works on human-AI thought partnership and eliciting richer human feedback
Data Workers' Inquiry (2024–25); TIME100 AI 2025 — labour conditions & power in data annotation
RewardBench (2024), Tülu 3 post-training recipes, RLHF Book — public reference on preference-data pipelines & costs
Visual Genome crowdsourcing; human-centred multimodal data & evaluation (e.g. DOCCI / dense annotation work) at UW/AI2
Rico (UIST 2017): 66K-screen mobile UI dataset underlying most UI-understanding/generation work
Senior author 'Artificial Artificial Artificial Intelligence' (2023) — crowd workers widely use LLMs; data-quality threat to human data
Consent in Crisis (NeurIPS 2024) and Data Provenance audits — licensing/supply shock for training data, a driver of paid human data
AlpacaFarm (2023) & AlpacaEval: simulating/replacing human preference feedback — the 'cheap synthetic vs. expensive human' question
First author 'Artificial Artificial Artificial Intelligence' (2023) and 'Prevalence and Prevention of LLM Use in Crowd Work' (CACM 2025)
'En attendant les robots' (2019) and DiPLab studies of micro-work and AI data supply chains (Madagascar, Latin America)
Data Shapley (ICML 2019) — data valuation, i.e. pricing individual training examples
Fairwork cloudwork ratings of annotation platforms; 'Feeding the Machine' (2024)
'Ghost Work' (2019, with Siddharth Suri) — canonical account of on-demand data labour
Open Problems and Fundamental Limitations of RLHF (2023) — catalogue of human-feedback data-quality failure modes
Public professional sources, every entry linked
Model technical reports and their credits, papers, job posts, company blogs, vendor case studies, talks. Roles are checked against 2025–26 evidence because this market reshuffles fast; the evidence date sits on every line. The structure is in BUYERS.md.
Almost no lab names its head of human data
Budget owners are the hardest role to find in public — most rosters stop at research leads and open job posts. Closing that gap is the first job of pass 2.