Miju Labs

The security dossier

The taste read

The lane is chosen, so the question is what is true about it. Mercor already sells the premium version at $150–250/hr, the arena layer holds 6,047,075 image votes collected free, the data costs $0.91 a comparison, and a 10–25x pricing contradiction sits unresolved at the centre. Three things are genuinely in the client's favour and one of them is a property right no US competitor can hold.

medium confidence12 minupdated 2026-08-31verdict · scoring · oracle · arenas · database right · pricing contradiction

Build the panel, not the corpus. The defensible business in taste is a calibrated, named, consented professional panel and the sealed evaluation content that falls out of it — sold as an instrument with a published reliability figure attached, never as a file of preference pairs. That shape is argued at What to build first. This page is why nothing else in the lane survives the evidence, and why the two moves everyone reaches for first — volume, and a public benchmark — are gone.

The decision to enter has been taken, so this is not a page about whether. It is about what the entry looks like from inside: worse on four dimensions than the pitch, better on three.

The five facts that define the entry

One: Mercor already sells this, at a higher price, from better pedigree. Its Senior Design Expert posting pays $150–250/hr to capture "how expert designers think: how you approach a problem, what separates strong craft from weak", in sessions recorded remotely for an unnamed client (Mercor). Its Agency Brand Design Expert posting pays $80–150/hr, requires agency pedigree "(e.g. Pentagram, Wolff Olins, Landor, Collins, IDEO)", and asks the expert to "design precise, task-specific grading criteria" and "score AI-generated and human work samples against those criteria, with detailed written justifications", open to "peer calibration" (Mercor).

That is rubric authoring, scored critique with rationale, peer calibration and narrated reasoning capture — the entire premium half of the artefact taxonomy — bought today by the largest company in expert data. Contra Labs' published "up to $100/hr" is not the top of the market; against a $65/hr freelance branding rate and a $65/hr median across all AI-training work it is the middle of the band. The only seam is that Mercor runs its creative vertical with generalist project leads and no design staff: it can source the designers but not design the study (The generalists in creative).

Two: volume is permanently lost, and it was lost for free. Arena's boards carry 6,047,075 image votes across 76 models and 645,773 video votes across 46, as of 25 August 2026 (arena.ai; arena.ai). Contra's flagship benchmark is 5,940 pairwise judgements from about thirty evaluators, and its Design Crit set 14,400 comparisons from ten designers (arXiv 2606.30561). Roughly 420x the specialist's entire published output, at a marginal cost of server time, from a company doing $100M annualised on 28 people and retaining 80% of the corpus (TechCrunch; Contrary).

The only four things left to sell

A vote is a click, and a click carries no identity, no reason, no axis and no process. The specialist can therefore compete on exactly four things — rater identity, written rationale, multi-axis rubrics, and captured process — and on nothing else. Any line that reduces to "we also have preference pairs" is priced against zero. The arena layer has the arithmetic.

Three: the dangerous competitor is Taste Labs, not Contra. Contra has the artefacts — Taste Labs has no dataset, no benchmark, no paper — and Taste Labs is still the more dangerous company. Founder Thais Castello Branco led growth at Exa, described by her pre-seed investor as "a Benchmark-backed AI startup serving foundation labs" (Latitud); she has sold to this buyer population and Contra's marketplace founders have not. Four of its eight open roles are ML or applied research; none of Contra's five are (tastelabs.com). Recruiting runs on a trust-propagation model where vetted evaluators nominate others (Amplify). And its published research agenda commits to design intent inference from edit sequences and design history beyond code diffs (tastelabs.com) — which is the trajectory asset, Contra's only real moat, announced months before anyone ships it. Workings at Taste Labs, in full.

Four: the data is cheap, and the competitor published the number. From the Contra × Lica TASTE paper: ten professional designers in two cohorts of five, "paid a flat project fee that averaged approximately $90 per hour", working ~13 and ~16 hours, producing 14,400 pairwise comparisons across nine criteria (arXiv 2605.20731).

The build-versus-buy arithmetic a lab does in one meeting

Ten designers × ~14.5 hours × $90 ≈ $13,050 for 14,400 comparisons ≈ $0.91 per comparison. Extrapolated: ImageReward's 137,000 is ~$125,000; HPD v2's 798,090 is ~$726,000; HPDv3's 1,170,000 comparisons is ~$1.06M of direct annotator labour. A frontier-competitive professional preference corpus costs about a million dollars — a fundable line item, not a moat. Arena's 6M-vote board would be $5.5M at the same rate, and Arena paid nothing.

The discipline that follows governs everything: the margin exists only on high-touch, low-volume, high-ASP work — calibration, rubric design, adjudication, narrated process, sealed suites. Never on comparison count.

Five: a 10–25x pricing contradiction sits at the centre and nobody has crossed it. Published practice pays $2.80/hr for HPS v2's fifty labellers and seven checkers, $12/hr for GenAI-Bench annotators, $16/hr for HEIM's crowdworkers, and $0 for UI-Bench's 194 hand-selected experts, thanked for "generously contributing their time and judgment" (arXiv 2306.09341; arXiv 2406.13743; arXiv 2311.04287; arXiv 2508.20410). The specialists quote $50–250/hr (Fast Company).

No disclosed buyer has yet paid the second number for the first job. HEIM states its bill outright — "the total amount spent for human annotations was $13,433.55" — for a field-defining, multi-axis, aesthetics-inclusive benchmark. The business is a bet that the low number must rise to meet the high one, and the record contains no instance of it happening. Sizing the taste market bounds the consequence: $50–300M of external spend, central $100–150M, smaller than Arena's revenue alone.

The three things genuinely in the client's favour

The EU database right is an asymmetric property right. Article 7 of Directive 96/9/EC lets the maker of a database prevent extraction of a substantial part where there has been substantial investment in obtaining, verification or presentation. It is not copyright and does not care whether anything in the corpus is original — pure investment protection, exactly the shape of a corpus of expert judgements. Article 11 restricts it to EU nationals, residents, and companies formed under Member State law with a genuine link to a Member State's economy. Contra.Work Inc., Taste Labs, Intelligence, AfterQuery and Mercor are US companies. None can hold it over its own corpus.

The trap is real: British Horseracing Board v William Hill, C-203/02, excludes investment in creating data, and a commissioned corpus is created data (Fieldfisher). The response is an accounting decision taken in month one — ledger verification and presentation separately, and layer a curated "obtained" tranche of award archives alongside the commissioned one. Everything the manufactured oracle already requires is verification spend; the question is whether the books prove it. The one right a US competitor cannot hold covers the structure, including that the EU entity must be the maker and not a subsidiary of a Delaware parent.

The supply-side opening is larger than the sentiment surveys imply, and unclaimed. The Sora revolt is the closest thing to a controlled experiment anyone has run: roughly 380 verified artists leaked alpha access, writing that they had been "lured into 'art washing'", objecting to unpaid feedback labour, to being used for PR, and to the requirement that "every output needs to be approved by the OpenAI team before sharing" (Newsweek). Read what is absent: nobody said evaluating AI output was beneath them. Every demand was a term sheet. The 48,000-signature Statement on AI Training turns on the word "unlicensed" and is textually about training works (ed.newtonrex.com). Cosmos gave creatives at Nike, Apple and Amazon a free toggle to block AI imagery and 10% used it (Fast Company).

And the largest unclaimed trust differentiator is in plain sight: not one platform in this market names its client. Contra, Mercor, Outlier, Taste Labs and Handshake all sell "work for frontier AI labs" without saying which — while 50.97% of surveyed artists "need no payment but object to who profits" (arXiv 2401.15497). And the incumbents wrote the counter-pitch themselves: Adobe's Firefly bonus is paid "at Adobe's discretion" in an undisclosed amount, with "no opt-out option for Adobe Stock content" (Adobe Help). See Will they say yes.

The headroom in taste is enormous and independently measured. On the Visual Aesthetic Benchmark — 400 tasks, 1,195 images, consensus from ten independent expert judges per task — the best frontier system identifies both best and worst image correctly in 26.5% of tasks, against 68.9% for human experts (arXiv 2605.12684). On Design Crit, off-the-shelf VLM judges agree with designers about 54% of the time against a 50% chance floor, a small trained head reaches 61.1%, and the human ceiling is 74.1% (Design Crit). Two groups, different formats, same conclusion — while the objective benchmarks saturated so hard they now mismeasure, GenEval having drifted to 17.7% absolute error before being replaced (arXiv 2512.16853). Where the headroom is has the landscape.

The strongest argument against, at full volume

Pick-a-Pic's own conclusion is that its results "showcase the importance of using real users as ground truth rather than expert annotators when collecting human preferences" — expert baseline 68.0%, PickScore 70.5% (arXiv 2305.01569). Design Arena is that argument commercialised: 5.3M users voting for nothing. And Luma is hiring a Research Engineer for "turning human evaluation frameworks into automated ones and wiring the results straight back into training" (Ashby). Taken together: the experts are not better, the crowd is free, and the buyer intends to automate you.

It measures a different thing than it is quoted for. PickScore was trained on 500K+ user preferences and evaluated on that same distribution, and its "expert" baseline was not a calibrated panel — the paper's experts are "friends and colleagues of the authors." The finding is that a model trained on user preference beats a handful of unscreened acquaintances at predicting user preference. It is silent on whether a screened panel predicts professional acceptability, and the literature runs the other way: HPSv3 moved to professional artists screened at ≥16 of 20 with 9–19 annotators per pair, raising convergence from 59.9% to 76.5% (arXiv 2508.03789).

Concede the domain split rather than argue with it. Where the target is consumer preference — which image an ordinary person likes — the user is the ground truth by definition, the arena is right, and the specialist should stay out. Where the target is professional acceptability — will an art director ship this, does the baseline grid hold at the third module, does the hierarchy invert — the user is structurally not the ground truth, because the buyer is selling into a professional workflow rather than to the user's taste. That is why Canva's Principal Research Scientist posting says design quality "gates everything else", and why What to build first runs in product and UI design rather than consumer image aesthetics.

And Luma's posting is a purchase order, not a threat. You cannot turn human evaluation frameworks into automated ones without human evaluation frameworks, and the automated metric then drifts — which is what GenEval's 17.7% and HPSv3++ needing 212K fresh pairs to re-condition on current-generation outputs both demonstrate (arXiv 2606.14657). Automation converts a one-off labour line into a recurring calibration line, which is the better business. What none of that fixes: if the client cannot say in one sentence why a lab building a consumer image model should pay a hundred times more for a professional's opinion of an image consumers will judge, the objection wins in that segment, and it should. The answer is segment selection, not rhetoric.

The six scores

AxisTasteSecurityHealth
proof344
budget354
defense334
cost524
reach545
room223
total212024

proof: 3. The scoreboards are public, the buyers are visibly losing on them, and publishing is cheap. But proof means a reference customer who says so out loud, and no company in this vertical has ever got a buyer to name it. Worse, the obvious credibility play is taken: AfterQuery built UI-Bench — the most rigorous public human-judged design benchmark — roughly a year before Contra launched, from 194 hand-picked experts at apparently zero cash cost. Security scores 4 because UK AISI publishes what it paid its suppliers; nothing equivalent exists here. What is genuinely unclaimed is narrower and better: nobody has published the panel-size-to-reliability curve for professional design evaluation, and a method is harder to copy than a leaderboard.

budget: 3. Money moves and it is legible: Mercor's live bands for a frontier client; OpenAI's Human Data team naming Sora at $207K–$385K (Ashby); Canva's Principal Research Scientist explicitly allocating "human data spend", the only named budget line at any app-layer company; Google's Imagen 3 report disclosing 366,569 ratings from 3,225 raters across 71 nationalities (arXiv 2408.07009). Against it: "aesthetic" appears in none of OpenAI's 753 postings; no posting across 1,235 model-builder listings names an external vendor; and the app layer shows one evaluation role in 1,562 postings, with $31.6B of vibe-coding valuation employing nobody to evaluate design quality (The creative tool layer). No Astra event, no contract value anywhere.

defense: 3, and the halves point opposite ways. In favour: the content decays by construction, because preference data is indexed to a model generation — a subscription written into the physics — and the manufactured oracle cannot be rebuilt from a delivered file, because a buyer can reconstruct the labels but not the panel's identity, the calibration history or the reliability record. Against: the supply is abundant and getting cheaper. About 1.06M US professionals across eleven segments, BLS projecting graphic design at −2% and naming AI as the cause, freelancers billing 22.4 of 40 hours so the marginal hour competes with an empty calendar, and a clearing price of $60–85/hr with no premium paid for creative expertise at all (Eleven labour markets, not one). Health scores 4 because the state maintains the credential. Here the only scarcity you can own is one you construct.

cost: 5 — the cheapest entry in the atlas. The competitor's entire published corpus cost under $20,000 of designer time; a flagship-scale study is $30–80K of direct expert cost; Sidebar sells a slot at $950 for 2,000–4,000 clicks, $0.32 a click, the only published cost-to-reach in this atlas (Sidebar). No rigs, no licences, no clinical estate, no PHI, no SOC 2 gate before the first pilot. Pushing back: panel depth multiplies everything, and trajectories cost $400–1,500 against $0.55 for a synthetic one (arXiv 2412.09605). Held at 5 because the first defensible artefact still costs half its equivalent in health.

reach: 5, and it is the best-argued 5 on the site. Every year the award bodies publish, by name and country, the people their own field selected as fit to judge: D&AD 300+ judges from 47 countries, ADC ~270 from 48, the Motion Awards 200+, Awwwards a minimum of 18 jurors per site (D&AD; The One Club; Awwwards). That is 800–1,000 named, peer-selected senior creatives a year, free to identify, and ADC's discipline taxonomy is the segment taxonomy. Below it sits ADPList's 38,024 mentors with public employer-stated profiles (ADPList). The caveat is priced in: 80% of creative professionals entered no award last year, so the registers are the seed and never the funnel (Creative Boom). Where the good ones actually are has the scorecard.

room: 2, and this is the binding constraint. In five months the lane went from empty to three funded direct entrants, one acquisition and one corpse: Taste Labs at $18.5M co-led by CRV and Amplify, Design Arena's parent at $7.9M led by Index, Lica acquired by Gamma on 25 August 2026 taking Contra's only ML partner with it, and Yupp dead on 15 April 2026 after raising $33M from a16z, having reached 1.3M signups and millions of preferences monthly without finding a sustainable buyer (The AI Cemetery). Off a 1 because nobody occupies the calibrated-named-consented-panel position, nobody sells the divergence map as a SKU, nobody has published per-axis reliability with variable panel depth, and no US competitor can hold the database right.

21, against security's 20 and health's 24. What the comparison means, now the lane is chosen, is narrow: expect security's economics, not health's — a crowded market with a small pot, where the position must be constructed rather than occupied, and where the two weakest axes are precisely the two that respond to construction rather than to effort. What it does not mean is that the decision should be revisited. The composite sorts; it does not decide. Its remaining job is to say where the money goes: into the panel and the protocol, which is all of defense, and into the position nobody has taken, which is all of room.

When the answer flips

To no. If the 300-respondent survey at the top of What the taste dossier could not establish shows creatives price selling judgement the same as selling work, the load-bearing supply assumption is inference all the way down. If Mercor posts a design-methodology or rubric-research role, the seam closes and room goes to 1. If a lab discloses an aesthetic campaign at crowd volume and crowd price, the 10–25x contradiction resolves downward and this is a low-cost-geography business. If the ablation in Ninety days in taste shows narrated trajectories are indistinguishable from $0.55 synthetic ones, the most defensible unit is not defensible. And if Adobe points Behance's 50M+ members at a paid evaluation programme — existing payments relationship, a model to improve, a demonstrated willingness to use contributor work without opt-out — the supply advantage evaporates in a week.

To a much louder yes. If any buyer discloses a price for panel evaluation above roughly $100/hr, the unit economics close on public information for the first time; not one rate card exists in this market today, from any company, for any artefact. If a lab's Article 53 training-data summary names a commissioned licensed corpus — enforcement powers went live on 2 August 2026 with fines to 3% of global turnover, and Anthropic, Mistral and xAI substituted prose for the template (Sell the paperwork with the data) — provenance converts from a nice-to-have into a procurement requirement. If the panel-size study shows required depth varies by more than a factor of two across axes, the pricing schedule in Phase decomposition is real and ownable. And if the Third Circuit in Thomson Reuters v. Ross accepts lost licensing to other AI companies as market harm, the licensed market that theory presupposes is this one (Copyright is the weakest thing you own).

The shape that survives is at What to build first, the sequence for testing it at Ninety days in taste, the gaps at What the taste dossier could not establish, and the people who can close them at Seventy people and one missing bridge and [/people]. Same rubric, other lanes: The security read, The health read, The specialist wedge, Expert data for frontier labs.