Miju Labs

The security dossier

The generative-media labs

One lab has a disclosed procurement function and it covers Sora — but the word 'aesthetic' appears in none of its 753 postings. Google discloses 366,569 ratings and never a price. And across the whole literature the same split recurs: designers write the prompts, a general crowd supplies the judgement.

high confidence10 minupdated 2026-08-30openai · deepmind · black forest labs · kling · luma · buyers · procurement

The prestige buyer for creative expert data is the group of labs that build image, video and audio models. The evidence about how they actually source aesthetic preference is better than expected in three places and completely absent in two, and the pattern that connects them is the single most useful thing on this page.

Across the model builders with machine-readable job boards — OpenAI (753 postings), ElevenLabs (248), Higgsfield (69), Suno (62), Luma (48), Black Forest Labs (16), Krea (10), Pika (10), Hedra (8), Recraft (4), Stability (4), Ideogram (3) — only four postings out of 1,235 are dedicated human-data or evaluation procurement roles: three at OpenAI and one at Luma. Not one posting in the sector names an external data vendor. Full posting bodies were searched for Scale AI, Surge, Mercor, Handshake, Invisible, Snorkel, micro1, Pareto, Labelbox, Appen and Toloka; the only hits were false positives from boilerplate.

So what

Procurement relationships in this sector do not appear in job descriptions. That is a finding about method, not about demand: you cannot map this buyer set by scraping job boards, and a client who concludes "nobody is buying" from ATS silence has drawn the wrong inference from the right data.

OpenAI — the only disclosed procurement function, and it covers Sora

OpenAI's Human Data team is the clearest buyer in the landscape and it names generative media explicitly:

In their words

"OpenAI's Human Data Team creates custom data solutions driving groundbreaking research. Our work enhances and evaluates our flagship models and products like ChatGPT, GPT-5, and Sora." — Research Program Manager, Human Data Campaigns

The mechanics are described plainly across three roles: "a key interface between our research roadmap, external vendors, AI trainers, and the Human Data engineering team"; "advise and empower program managers and vendors to drive day-to-day execution"; and from the Program Manager posting, "gather requirements, write instructions, define success criteria, and calibrate the AI trainers" (Ashby). The remit covers "bespoke data campaigns, scalable synthetic data generation, and product-embedded signals."

Bands of $207K–$385K across the three roles (Ashby) say this is a resourced function, not an experiment. The unit of purchase is a campaign. Anyone selling here is selling into a campaign structure with internal program management, internal platform engineering, and external supply — which is the shape Expert data for frontier labs describes for the whole sector.

The word that is missing

"Aesthetic" appears nowhere in OpenAI's 753 postings. No designer-specific sourcing could be established. The Sora 2 system card discloses only that "we worked with internal red teamers" — no external creative panel, no aesthetic evaluation methodology, no vendor, no counts (OpenAI). For a flagship video launch that is a striking non-disclosure, and whether large aesthetic campaigns exist behind it is the swing factor for the whole market estimate.

Google DeepMind — volume disclosed, price never

The Imagen 3 technical report is the most detailed procurement-adjacent document available anywhere in this field (arXiv 2408.07009):

In their words

"In total, we collected 366,569 ratings in 5,943 submissions from 3,225 different raters. Each rater participated in at most 10% of our studies, and in each study, each rater provided approximately 2% of the ratings, to avoid biasing the results to a particular set of raters' judgments. Raters from 71 different nationalities participated in our studies."

The evaluation aspects include "visual appeal" — an explicit aesthetic axis — measured by side-by-side comparison rather than absolute scoring, because "quantitative judgment (e.g. assigning a score between 1 and 5) is in practice difficult to calibrate across raters." That methodological note is worth keeping: it is the same reason serious panels use pairwise designs, and it is a fact the client can cite when a buyer asks why not just score things.

On external work, the report confirms Google pays: "the groups testing Imagen 3 were compensated for their time. External groups design their own methodology... Reports are written independently of Google DeepMind." But those groups are safety and CBRN red teams from academia, civil society and commercial organisations — not creative professionals.

Gap in the record

Rater pay, the vendor supplying the 3,225 raters, and any spend figure. Google discloses volume and never price, in the most transparent document the field has. No Veo 3 or Gemini-image technical report with comparable human-evaluation methodology was retrievable.

The pattern that runs through everything: designers write, crowds judge

Now the connection that matters more than any individual lab.

Imagen 3's prompt set was GenAI-Bench: "a set of 1600 high-quality prompts collected from professional designers" (arXiv 2406.13743). So Google bought designer-authored prompts, via an academic benchmark, and then used a general pool of 3,225 raters from 71 nationalities to supply the judgements.

That split is not a Google idiosyncrasy. It recurs across the whole record:

StudyProfessional inputJudgement input
Imagen 3GenAI-Bench prompts from professional designers3,225 general raters, 71 nationalities
GenAI-Benchprompts from professional designerscrowd annotators at ~$12/hr
HEIMtask and rubric designcrowdworkers, ≥5 per image, $16/hr
HPS v2screening test, 61% hire rate57 contractors — "average annotators rather than experts"
Pick-a-Picauthors' friends and colleagues as an expert baselinereal product users, unpaid
The central pattern of this market

Professional creative input is bought at the specification layer — prompts, rubrics, task design, screening criteria — and crowd labour is bought at the judgement layer. Every disclosed procurement in the record fits this shape. A vendor selling professional judgement at volume is selling into the half of the pipeline the buyers have consistently chosen not to professionalise; a vendor selling rubric and study design plus a small calibrated panel is selling into the half they already pay professionals for. That is the difference between a wedge and a wish, and it is the same conclusion The oracle problem reaches from the measurement side.

Black Forest Labs — less opaque than assumed

BFL is described everywhere as a black box. It is not entirely: it runs a live Ashby board with 16 postings, and its post-training role states the aesthetic training target outright:

In their words

"Advance techniques across the post-training stack: SFT, RLHF, RLAIF, DPO, preference learning, and reward modeling to align models with human intent and aesthetic judgment... Own the full post-training pipeline end to end — from data curation and reward modeling through fine-tuning, preference optimization, distillation, safety tuning, evaluation, and deployment." — MTS, Post Training, Freiburg

This is the clearest statement by any model builder that aesthetic judgement is a training target requiring preference data. The same role mentions "personalization and customization capabilities that let users adapt our models to their own creative style."

It proves demand. It says nothing about supply: BFL has published no technical report with human-evaluation methodology for FLUX, discloses no annotator counts, names no vendor, and runs no artist programme that could be found. One senior post-training hire in Freiburg is the entire visible surface — and, for a European operator, the most reachable one on this page.

The Chinese labs are the most forthcoming about using professionals

Counter-intuitively, the labs that say least about pay say most about credentials.

Kuaishou (Kling). The Kling-Omni technical report states that "to bridge the gap between model outputs and human aesthetic preferences, we employ Direct Preference Optimization (DPO)"; that the team "engaged with creators ranging from professional directors to general users" to construct the evaluation system; and that it ran "a double-blind human evaluation... inviting domain experts and professional annotators to compare Kling-Omni against industry-leading models" (arXiv 2512.16776). Counts, pay and vendor: not disclosed.

That is the strongest published evidence anywhere that a lab uses professionals for both requirements-setting and evaluation. It is also the client's best citation, and it is Chinese — which matters for The one right a US competitor cannot hold and for anyone whose go-to-market assumes American buyers move first.

ByteDance (Seedream). Seedream 3.0 reports "a systematic evaluation by human experts" on a 377-prompt benchmark across text-image alignment, structural correctness and aesthetic quality, plus a defect detector trained on 15,000 manually annotated samples and an "Aesthetic Caption" post-training stage (arXiv 2504.11346). No counts, no pay, no vendor.

Alibaba (Wan) is the counter-example and deserves its weight. Wan's published evaluation is almost entirely automated: aesthetics scored by models, stylisation by Qwen2-VL, camera movement by RAFT optical flow, and data curation by "expert-based collection" in which an expert model selects the top 20% of images (arXiv 2503.20314). Human aesthetic panels do not feature in the methodology at all. Wan is evidence for the automation thesis, not the human-data thesis, and a serious buyer conversation will raise it.

Luma states the automation counter-thesis explicitly

Luma runs the only other dedicated evaluation role in the sector, and it cuts directly against the client:

In their words

"You'll be turning human evaluation frameworks into automated ones and wiring the results straight back into training... Work with researchers running human studies to translate their frameworks into automated or semi-automated systems." — Research Engineer, Evaluations, $190K–$375K (Ashby)

Read honestly, that is a company treating human studies as seed corn for automated metrics, not as a recurring purchase. It is the single most important sentence for anyone modelling this as subscription revenue. The optimistic reading — that every automated metric needs periodic human recalibration, and drift makes that recurring — is plausible and unproven; the pessimistic reading is that a lab buys the panel once, distils it into a reward model, and never returns. The oracle problem is where that argument is settled, or at least stated properly.

Luma's other creative hiring is employment rather than procurement: Creative Technologists in London and LA at $145K–$205K, Forward Deployed Creatives at the same band, product and visual designers at $225K–$300K.

Everyone else, briefly

Suno gives the cleanest verified market rate for contract creative labour in this dataset — $75–$90/hour for contract designers, $80–$120/hour for a contract creative strategist — plus a Director, Editorial & Curation at $200K–$263K seeking "a trusted taste-maker". That is editorial curation for the product, not training data (Ashby).

Higgsfield answers "where do we get creative judgement" with a low-cost in-house studio: its Concept/VisDev Artist, Film Director, Motion Designer and Product Designer roles are all in Almaty, Kazakhstan.

Krea has raised over $83M and built Krea 2 "for aesthetic diversity and stylistic control" with an in-house team of "musicians, designers, visual artists, and engineers." Recraft runs an AI Designer applying the model to "real creative scenarios to uncover where it fails" — an in-house aesthetic evaluator by another name, and precisely the function a vendor could displace. Stability's Research Scientist, Professional Creative Workflows is asked to "design datasets, benchmarks, and evaluation protocols for professional workflows where quality, editability, consistency, brand adherence, and production usefulness matter" (Greenhouse) — a build-it-yourself mandate describing the client's product almost exactly.

Runway's Hundred Film Fund is a $5M total fund granting up to $1M per film — a content and marketing subsidy, not data procurement.

A research trap

jobs.ashbyhq.com/runway is a business-planning startup, not RunwayML"Runway replaces traditional spreadsheets with a modern planning platform." RunwayML has no machine-readable board. Any competitive analysis citing "Runway" ATS data should be checked for this confusion before it reaches a slide.

Midjourney is completely dark to this method. No public applicant-tracking system (Greenhouse, Ashby and Lever all 404), no technical report, no model card, no published evaluation methodology, and nothing retrievable about human raters. What is known is structural rather than documented — its products have historically surfaced grid choices that produce preference data as a by-product — and that is not enough to assert anything. A company with no recruiting surface and no published methodology is not running a procurement process, so it is not a realistic near-term buyer regardless of how much preference data it holds.