Miju Labs

The security dossier

Sizing the taste market

$50–300M of externally-purchased spend, central $100–150M, built from a $6–10B expert-data market with a $3.8B verified floor times a 1–3% creative share. The uncomfortable half: published practice pays $0–16/hour for aesthetic judgement, and a field-defining benchmark cost $13,433.55.

low confidence8 minupdated 2026-08-30market size · pricing · gross vs net · heim · genai-bench · estimates

The specific market — human aesthetic preference and creative evaluation sold as a service in 2026 — cannot be sized from public information. No vendor breaks out creative revenue. No lab discloses spend. No analyst series covers the vertical. Anyone quoting a figure is extrapolating, including this page.

What can be done is to bound the containing market, apply a defensible share, and cross-check from three other directions. Here is the working, with every input visible so a reader can disagree with a specific step rather than with the conclusion.

Step one: the containing market has a verified floor

VendorFigureGross or netSource
Mercor$614M H1 2026 revenue; ~$2B run rate (June); $2.8B forecast year-end; 91% from foundation-model companies; $20B valuationGROSS — this is marketplace volume, of which Mercor keeps roughly a thirdBigGo Finance
Surge AI$1.2B revenue, bootstrappedreported as revenue; take-rate structure not disclosedSacra [WEAK]
micro1$500M gross run rate (Aug 2026), from $7M a year earlierGROSS, stated as suchTechCrunch
AfterQuery$100M+ revenue run rate; $30M Series A at $300Mself-reported on its own siteafterquery.com
The floor and the frame

Four vendors alone disclose roughly $3.8B of run-rate revenue. Adding Scale, Invisible, Turing, Handshake, Snorkel, Labelbox, Appen and internal lab spend, a $6–10B human-expert-data market for 2026 is defensible. Note that the largest single input is gross marketplace volume, not net revenueMercor keeps roughly a third of what passes through it, so apply GMV is not revenue before treating $3.8B as an addressable pot. Surge and micro1 disclose no take-rate structure at all. Sector commentary about a "near-$100 billion" business refers to valuations, not revenue, and must not be used here.

Mercor names its customers — OpenAI, Anthropic, Google DeepMind, Reflection AI and Thinking Machines Lab, with Meta suspended after a March security incident — which is what makes the 91%-from-labs figure usable. Expert data for frontier labs carries the full picture of that buyer set.

Step two: the creative share is 1–3%, and the justification is an absence

Mercor's disclosed domains are "lawyers, doctors, writers, and researchers." No design or visual-arts vertical is named at any major vendor. Micro1, Surge and Handshake coverage foregrounds law, medicine, finance and software. Handshake's top rates — $300–500/hr — go to physicians, STEM researchers and radiologists, and it has no design vertical at all.

Creative work is plainly being done — Mercor runs a creative vertical with live design postings, as The generalists in creative documents — but it is not large enough for any vendor to name. That supports a share of 1–3%, which against a $6–10B market gives $60M–$300M.

Step three: three cross-checks

Bottom-up from the labs. The realistic buyers of aesthetic preference data number roughly 17: OpenAI, Google DeepMind, Black Forest Labs, ByteDance, Kuaishou, Alibaba, MiniMax, Midjourney, Runway, Luma, Ideogram, Recraft, Krea, Stability, Suno, ElevenLabs, Udio. Imagen 3's disclosed campaign was 366,569 ratings from 3,225 raters (arXiv 2408.07009). At crowd rates of $12–16/hour and roughly 100 ratings an hour, that campaign costs about $44K–$59K. Twenty such campaigns a year across 17 organisations is $15–20M a year at crowd prices. At professional rates of $50–250/hour the same volume is $60M–$400M — but only under total conversion to professional panels, which nothing in the record suggests is happening.

Bottom-up from the app layer. Canva is the only app-layer company with a confirmed "human data spend" line, and its supply is largely in-house teams in the Philippines and the EU (The creative tool layer). If Canva spends single-digit millions and is the most advanced buyer in its cohort, the whole layer plausibly spends $10–50M/year, mostly on salaries. The externally addressable portion today is likely under $10M.

From the vendors selling it. The three companies explicitly selling creative taste have raised $7.9M (Intelligence), $18.5M (Taste Labs) and an undisclosed amount inside a $45.2M parent (Contra Labs). Seed-stage vendors with no disclosed customers do not indicate a large realised market. AfterQuery's $100M+ is generalist expert data, and its design work was a free-panel marketing exercise.

The answer

Read

Externally-purchased human aesthetic preference and creative evaluation data, globally, 2026: $50M–$300M, weighted toward the lower half. Central estimate $100–150M.

Working: a $6–10B expert-data market (verified floor $3.8B, largely gross) × a 1–3% creative share (justified because creative is not a named vertical at any major vendor) = $60–300M; cross-checked bottom-up at $15–20M for lab campaigns at crowd prices; cross-checked against an app layer spending under $10M externally; cross-checked against vendors who have collectively raised under $30M in the category. Confidence: low-to-moderate. Every input except vendor revenue is inferred.

For scale: that whole market is smaller than Arena's single-company annualised revenue of $100M, and roughly one twentieth of Mercor's gross run rate. A vendor taking 20% of the central estimate is a $20–30M revenue business — a real company, and not the one the "Mercor for design" framing implies. How much money is actually in the buyer pool and The taste read carry what follows from that.

The pricing contradiction at the centre of the business case

SourceRate for aesthetic or creative judgement
HPS v2 labelers (published)$2.80/hour — 20 CNY/hr labelers, 25 CNY/hr checkers (arXiv 2306.09341)
GenAI-Bench annotators (published)$12.00/hour (arXiv 2406.13743)
HEIM annotators (published)$16.00/hour (arXiv 2311.04287)
UI-Bench — 194 hand-selected design experts$0 (arXiv 2508.20410)
Suno contract designers (verified from ATS)$75–$90/hour
Mercor Senior Design Expert (company posting)$150–$250/hour
Contra Labs creatives on RLHF (reported)$50–$250/hour (Fast Company)
The 10–25× gap

Published buyer practice pays $0–$16/hour for aesthetic preference labour. The companies selling designer taste quote $50–$250/hour. No disclosed buyer has yet crossed that gap. The closest evidence in the record is Kuaishou engaging "professional directors" and "domain experts and professional annotators" for Kling-Omni evaluation (arXiv 2512.16776) — with price undisclosed. The client's entire business is the bet that the first number must rise to meet the second, and that bet is currently unwitnessed.

Two structural notes on why the gap has held. First, the professional input labs do buy sits at the specification layer — GenAI-Bench's prompts came from professional designers and were then judged by crowds, a split that recurs everywhere (The generative-media labs). Second, the highest-credential panel in the whole literature was also the cheapest, because UI-Bench's 194 experts were invited rather than hired. Expertise has been available for free to anyone with a plausible paper, which is a supply fact the pricing model has to survive.

The reference price for a field-defining benchmark is four figures

In their words

"We required five different annotators per sample... Based on an hourly wage of $16 per hour, each annotator was paid $0.02 for answering a single multiple-choice question. The total amount spent for human annotations was $13,433.55." — HEIM (arXiv 2311.04287)

HEIM evaluated dozens of text-to-image models across alignment, quality, aesthetics and originality. It cost thirteen thousand dollars. GenAI-Bench cost roughly $9,600 — about 800 annotator-hours at $12/hour — and Google DeepMind then used its designer-sourced prompts to evaluate Imagen 3. The single most consequential prompt set in commercial image evaluation cost less than a mid-market SaaS seat contract.

The same anchor from the professional side: the Contra × Lica TASTE study paid 10 designers a flat fee averaging ~$90/hour for 13–16 hours each — about $13,000 for 14,400 comparisons, or ~$0.91 each (arXiv 2605.20731). Extrapolated, a million-comparison frontier-grade corpus is ~$1M of direct annotator labour. That is a fundable line item for any lab with a program manager, not a moat — which is why the sellable unit has to be the thing that does not scale linearly: rubric design, calibration, named panels, narrated process. Eight units and one cost anchor and The trajectory moat carry that argument.

The swing factor, and it could move the estimate by an order of magnitude

OpenAI's aesthetic campaign volume for Sora is entirely undisclosed. The Human Data team explicitly covers Sora, runs "bespoke data campaigns", and manages external vendors at $207K–$385K band levels — but the word "aesthetic" appears in none of OpenAI's 753 postings, and the Sora 2 system card mentions only internal red teamers. If aesthetic campaigns inside OpenAI, Google and ByteDance are large, the $100–150M central estimate is low by 10×. I could not establish this either way, and it is the number that matters most. Everything else on the open-questions list — Mercor's and Surge's creative client mix, Contra's real contract values, Google DeepMind's rater vendor — is second-order next to it. See What the taste dossier could not establish.