The prestige buyer for creative expert data is the group of labs that build image, video and audio models. The evidence about how they actually source aesthetic preference is better than expected in three places and completely absent in two, and the pattern that connects them is the single most useful thing on this page.
Across the model builders with machine-readable job boards — OpenAI (753 postings), ElevenLabs (248), Higgsfield (69), Suno (62), Luma (48), Black Forest Labs (16), Krea (10), Pika (10), Hedra (8), Recraft (4), Stability (4), Ideogram (3) — only four postings out of 1,235 are dedicated human-data or evaluation procurement roles: three at OpenAI and one at Luma. Not one posting in the sector names an external data vendor. Full posting bodies were searched for Scale AI, Surge, Mercor, Handshake, Invisible, Snorkel, micro1, Pareto, Labelbox, Appen and Toloka; the only hits were false positives from boilerplate.
Procurement relationships in this sector do not appear in job descriptions. That is a finding about method, not about demand: you cannot map this buyer set by scraping job boards, and a client who concludes "nobody is buying" from ATS silence has drawn the wrong inference from the right data.
OpenAI — the only disclosed procurement function, and it covers Sora
OpenAI's Human Data team is the clearest buyer in the landscape and it names generative media explicitly:
"OpenAI's Human Data Team creates custom data solutions driving groundbreaking research. Our work enhances and evaluates our flagship models and products like ChatGPT, GPT-5, and Sora." — Research Program Manager, Human Data Campaigns
The mechanics are described plainly across three roles: "a key interface between our research roadmap, external vendors, AI trainers, and the Human Data engineering team"; "advise and empower program managers and vendors to drive day-to-day execution"; and from the Program Manager posting, "gather requirements, write instructions, define success criteria, and calibrate the AI trainers" (Ashby). The remit covers "bespoke data campaigns, scalable synthetic data generation, and product-embedded signals."
Bands of $207K–$385K across the three roles (Ashby) say this is a resourced function, not an experiment. The unit of purchase is a campaign. Anyone selling here is selling into a campaign structure with internal program management, internal platform engineering, and external supply — which is the shape Expert data for frontier labs describes for the whole sector.
"Aesthetic" appears nowhere in OpenAI's 753 postings. No designer-specific sourcing could be established. The Sora 2 system card discloses only that "we worked with internal red teamers" — no external creative panel, no aesthetic evaluation methodology, no vendor, no counts (OpenAI). For a flagship video launch that is a striking non-disclosure, and whether large aesthetic campaigns exist behind it is the swing factor for the whole market estimate.
Google DeepMind — volume disclosed, price never
The Imagen 3 technical report is the most detailed procurement-adjacent document available anywhere in this field (arXiv 2408.07009):
"In total, we collected 366,569 ratings in 5,943 submissions from 3,225 different raters. Each rater participated in at most 10% of our studies, and in each study, each rater provided approximately 2% of the ratings, to avoid biasing the results to a particular set of raters' judgments. Raters from 71 different nationalities participated in our studies."
The evaluation aspects include "visual appeal" — an explicit aesthetic axis — measured by side-by-side comparison rather than absolute scoring, because "quantitative judgment (e.g. assigning a score between 1 and 5) is in practice difficult to calibrate across raters." That methodological note is worth keeping: it is the same reason serious panels use pairwise designs, and it is a fact the client can cite when a buyer asks why not just score things.
On external work, the report confirms Google pays: "the groups testing Imagen 3 were compensated for their time. External groups design their own methodology... Reports are written independently of Google DeepMind." But those groups are safety and CBRN red teams from academia, civil society and commercial organisations — not creative professionals.
Rater pay, the vendor supplying the 3,225 raters, and any spend figure. Google discloses volume and never price, in the most transparent document the field has. No Veo 3 or Gemini-image technical report with comparable human-evaluation methodology was retrievable.
The pattern that runs through everything: designers write, crowds judge
Now the connection that matters more than any individual lab.
Imagen 3's prompt set was GenAI-Bench: "a set of 1600 high-quality prompts collected from professional designers" (arXiv 2406.13743). So Google bought designer-authored prompts, via an academic benchmark, and then used a general pool of 3,225 raters from 71 nationalities to supply the judgements.
That split is not a Google idiosyncrasy. It recurs across the whole record:
| Study | Professional input | Judgement input |
|---|---|---|
| Imagen 3 | GenAI-Bench prompts from professional designers | 3,225 general raters, 71 nationalities |
| GenAI-Bench | prompts from professional designers | crowd annotators at ~$12/hr |
| HEIM | task and rubric design | crowdworkers, ≥5 per image, $16/hr |
| HPS v2 | screening test, 61% hire rate | 57 contractors — "average annotators rather than experts" |
| Pick-a-Pic | authors' friends and colleagues as an expert baseline | real product users, unpaid |
Professional creative input is bought at the specification layer — prompts, rubrics, task design, screening criteria — and crowd labour is bought at the judgement layer. Every disclosed procurement in the record fits this shape. A vendor selling professional judgement at volume is selling into the half of the pipeline the buyers have consistently chosen not to professionalise; a vendor selling rubric and study design plus a small calibrated panel is selling into the half they already pay professionals for. That is the difference between a wedge and a wish, and it is the same conclusion The oracle problem reaches from the measurement side.
Black Forest Labs — less opaque than assumed
BFL is described everywhere as a black box. It is not entirely: it runs a live Ashby board with 16 postings, and its post-training role states the aesthetic training target outright:
"Advance techniques across the post-training stack: SFT, RLHF, RLAIF, DPO, preference learning, and reward modeling to align models with human intent and aesthetic judgment... Own the full post-training pipeline end to end — from data curation and reward modeling through fine-tuning, preference optimization, distillation, safety tuning, evaluation, and deployment." — MTS, Post Training, Freiburg
This is the clearest statement by any model builder that aesthetic judgement is a training target requiring preference data. The same role mentions "personalization and customization capabilities that let users adapt our models to their own creative style."
It proves demand. It says nothing about supply: BFL has published no technical report with human-evaluation methodology for FLUX, discloses no annotator counts, names no vendor, and runs no artist programme that could be found. One senior post-training hire in Freiburg is the entire visible surface — and, for a European operator, the most reachable one on this page.
The Chinese labs are the most forthcoming about using professionals
Counter-intuitively, the labs that say least about pay say most about credentials.
Kuaishou (Kling). The Kling-Omni technical report states that "to bridge the gap between model outputs and human aesthetic preferences, we employ Direct Preference Optimization (DPO)"; that the team "engaged with creators ranging from professional directors to general users" to construct the evaluation system; and that it ran "a double-blind human evaluation... inviting domain experts and professional annotators to compare Kling-Omni against industry-leading models" (arXiv 2512.16776). Counts, pay and vendor: not disclosed.
That is the strongest published evidence anywhere that a lab uses professionals for both requirements-setting and evaluation. It is also the client's best citation, and it is Chinese — which matters for The one right a US competitor cannot hold and for anyone whose go-to-market assumes American buyers move first.
ByteDance (Seedream). Seedream 3.0 reports "a systematic evaluation by human experts" on a 377-prompt benchmark across text-image alignment, structural correctness and aesthetic quality, plus a defect detector trained on 15,000 manually annotated samples and an "Aesthetic Caption" post-training stage (arXiv 2504.11346). No counts, no pay, no vendor.
Alibaba (Wan) is the counter-example and deserves its weight. Wan's published evaluation is almost entirely automated: aesthetics scored by models, stylisation by Qwen2-VL, camera movement by RAFT optical flow, and data curation by "expert-based collection" in which an expert model selects the top 20% of images (arXiv 2503.20314). Human aesthetic panels do not feature in the methodology at all. Wan is evidence for the automation thesis, not the human-data thesis, and a serious buyer conversation will raise it.
Luma states the automation counter-thesis explicitly
Luma runs the only other dedicated evaluation role in the sector, and it cuts directly against the client:
"You'll be turning human evaluation frameworks into automated ones and wiring the results straight back into training... Work with researchers running human studies to translate their frameworks into automated or semi-automated systems." — Research Engineer, Evaluations, $190K–$375K (Ashby)
Read honestly, that is a company treating human studies as seed corn for automated metrics, not as a recurring purchase. It is the single most important sentence for anyone modelling this as subscription revenue. The optimistic reading — that every automated metric needs periodic human recalibration, and drift makes that recurring — is plausible and unproven; the pessimistic reading is that a lab buys the panel once, distils it into a reward model, and never returns. The oracle problem is where that argument is settled, or at least stated properly.
Luma's other creative hiring is employment rather than procurement: Creative Technologists in London and LA at $145K–$205K, Forward Deployed Creatives at the same band, product and visual designers at $225K–$300K.
Everyone else, briefly
Suno gives the cleanest verified market rate for contract creative labour in this dataset — $75–$90/hour for contract designers, $80–$120/hour for a contract creative strategist — plus a Director, Editorial & Curation at $200K–$263K seeking "a trusted taste-maker". That is editorial curation for the product, not training data (Ashby).
Higgsfield answers "where do we get creative judgement" with a low-cost in-house studio: its Concept/VisDev Artist, Film Director, Motion Designer and Product Designer roles are all in Almaty, Kazakhstan.
Krea has raised over $83M and built Krea 2 "for aesthetic diversity and stylistic control" with an in-house team of "musicians, designers, visual artists, and engineers." Recraft runs an AI Designer applying the model to "real creative scenarios to uncover where it fails" — an in-house aesthetic evaluator by another name, and precisely the function a vendor could displace. Stability's Research Scientist, Professional Creative Workflows is asked to "design datasets, benchmarks, and evaluation protocols for professional workflows where quality, editability, consistency, brand adherence, and production usefulness matter" (Greenhouse) — a build-it-yourself mandate describing the client's product almost exactly.
Runway's Hundred Film Fund is a $5M total fund granting up to $1M per film — a content and marketing subsidy, not data procurement.
jobs.ashbyhq.com/runway is a business-planning startup, not RunwayML — "Runway replaces traditional spreadsheets with a modern planning platform." RunwayML has no machine-readable board. Any competitive analysis citing "Runway" ATS data should be checked for this confusion before it reaches a slide.
Midjourney is completely dark to this method. No public applicant-tracking system (Greenhouse, Ashby and Lever all 404), no technical report, no model card, no published evaluation methodology, and nothing retrievable about human raters. What is known is structural rather than documented — its products have historically surfaced grid choices that produce preference data as a by-product — and that is not enough to assert anything. A company with no recruiting surface and no published methodology is not running a procurement process, so it is not a realistic near-term buyer regardless of how much preference data it holds.