Ranked by decision impact, with the action that settles each. Several are closable inside the first ninety days for less than the cost of one senior hire, which is the main reason to list them rather than caveat around them — and the first one is the assumption the verdict rests on.
1. Nobody has asked whether selling judgement differs from selling work
The whole business rests on this distinction and it has never been measured. The claim is that a designer who would refuse to license their portfolio for training will accept paid work evaluating model output — that the objection is to extraction rather than to participation.
The evidence for it is real but entirely indirect: the Sora protest letter, where every demand was a term sheet — unpaid, uncredited, approval-gated — and nobody said evaluating output was beneath them (Newsweek); the wording of the 48,000-signature Statement on AI Training, whose single sentence turns on the word "unlicensed" and is textually about training; and Cosmos's revealed 10% hard-refusal rate on a free AI-blocking toggle. Every one of those is an inference. No survey anywhere puts the two propositions side by side.
The action. A ~300-respondent survey, segmented by discipline and seniority, asking the same person to price both propositions. A few thousand dollars, three weeks, and it either confirms the central supply assumption or kills it before any money is spent on recruiting. It is also publishable, which makes it recruiting collateral as well as diligence. This is the highest-return research the client can commission and it is not close.
2. How many raters per axis, for what reliability
This sets the entire cost structure. Panel depth is the largest single driver of unit cost, and practice ranges across a factor of six with no published basis: Contra uses "3+ professional creative evaluators," the Visual Aesthetic Benchmark used 10 independent expert judges per task (arXiv 2605.12684), Awwwards runs a minimum of 18 jurors with automatic rejection of the three most-outlying scores (Awwwards).
No study establishes, for professional design evaluation, the panel size needed to reach a stated reliability target per axis. The two closest sources — the Konstanz expertise-screening work and the CrowdMOS methodology — were not retrievable during this research. And the requirement is certainly not uniform: agreement is highest on prompt adherence and lowest on visual appeal, and rises monotonically with workflow phase, so the same target reliability costs very different amounts on different axes (The oracle problem).
The action. Subsample your own pilot panel at n = 3, 5, 7, 10, 15 and plot achieved reliability per axis. It is a re-analysis, not a collection — one week of one person's time. It is publishable, it directly sets pricing, and no competitor has published it, which makes it the best available marketing artefact in a niche where the benchmark play is already taken.
3. No per-unit price exists in public for anything in this market
Every cost figure in this dossier is back-derived. The anchors are Contra's published expert rate of "up to $100/hr" with briefs "typically 5 to 10 hours," the throughput of roughly 15,000 professional judgements across 28 evaluators in the Human Creativity Benchmark, the $350 paid to each of 50 professionals in the formative survey, and the TASTE paper's disclosure of ten designers at approximately $90/hour producing 14,400 comparisons (which yields about $0.91 per pairwise comparison).
That is enough to build a cost model. It is not enough to build a price model, because no vendor publishes a per-unit price for any taste-data artefact and no buyer has disclosed one. The pricing pages of this market do not exist.
The action. Two things, neither of them research. Get a real quote from a competitor via a credible buyer-side introduction, and put a price in front of a buyer early enough that the answer is information rather than a lost deal. Until then, treat every margin figure on this site as an order-of-magnitude planning band and say so in any board paper.
4. OpenAI's aesthetic campaign volume for Sora
The swing factor on market size, and it is entirely undisclosed. OpenAI's Human Data team explicitly covers Sora and manages external vendors, but the word "aesthetic" appears nowhere in 753 OpenAI postings, and the Sora 2 system card discloses only that "we worked with internal red teamers" — no external creative panel, no methodology, no vendor, no counts (OpenAI).
The bottom-up estimate puts the addressable pot at roughly $15–20M a year at crowd rates across about 17 relevant labs, rising to $60M–$400M only if labs converted entirely to professional panels, which no evidence suggests they are doing. If large undisclosed aesthetic campaigns exist inside OpenAI, Google or ByteDance, that estimate is low by an order of magnitude.
The action. This is a conversation, not a search. Souki Mansoor runs the Sora artist programme and is publicly reachable; the diagnostic question is whether any evaluation signal from that programme fed back into training, because a yes means the budget exists and is simply not procured through a vendor yet.
5. Contra Labs' partner list
Contra's reported customer list — Framer, Webflow, Lovable, Replit, HeyGen — could be neither corroborated nor refuted, and contra.com/labs returns a 404. Relatedly, whether Contra's advertised $50–250/hour is a rate actually paid on real volume or an advertised ceiling is single-sourced.
This matters because it is the difference between a competitor with revenue and a competitor with a website. The action: ask the named companies directly rather than trying to source it. Product and design leads at Framer and Webflow are publicly reachable, and "are you buying design evaluation data from anyone" is a question people answer.
The same discipline applies to a second unverifiable claim about the leading investor-backed rival: Taste Labs' customers are not named anywhere, and the widely re-reported list of Runway, Figma, Character AI and Adobe appears to be a misreading of the investor announcement, which names them as companies facing the problem rather than as customers (Taste Labs, in full).
6. Design Arena's $60M ARR should not be used
Grace Li is quoted saying the site "is currently generating $60 million in ARR" (Yahoo Finance). It is self-reported, uncorroborated, and four things argue against taking it at face value: there is no consumer monetisation, so it is enterprise data-access revenue; comparable arena revenue is explicitly consumption-based and not recurring; the implied revenue per user is out of line with the much larger LMArena; and a company genuinely at $60M ARR does not raise a $7.9M seed (Design Arena, in full).
The action: none, and that is the point. Do not put this number in a fundraising deck, a market model or a competitive slide. If it is wrong, using it discredits everything next to it; if it is right, it is still not comparable to anything you will report. The generalisable rule for this whole category is that "ARR" is a category error — model it as project revenue with a re-run cadence.
7. There are no first-hand worker accounts of this job
Zero. Mercor's 669 Trustpilot reviews contain none from designers. There are none for Contra Labs, none for Taste Labs, none for Mercor's design experts. The experience record for this labour market is entirely unwritten.
That is a real risk and a real opportunity in the same fact. The risk is that nobody — including the client — knows what the work feels like after the fortieth brief, which is where retention is decided and where the reliability drift problem starts. The opportunity is that the client will be writing that record with its first hundred hires, and the first credible published account of the job, written by people who did it and were paid properly, is worth more to the supply side than any recruiting campaign.
The action. Instrument retention from the first cohort — completion rates, repeat-brief take-up, voluntary attrition, and an exit question — and treat the results as publishable rather than as internal ops data.
The shorter list
Real gaps, lower decision weight, grouped so they can be worked in a batch.
| Gap | Why it matters | How it closes |
|---|---|---|
| Aspen Hopkins's employment status at Contra | If she is an academic collaborator rather than staff, the competitor's academic credibility is rented and she is available | One email |
| Google DeepMind's rater vendor and pay | 366,569 ratings from 3,225 raters is a large managed panel and someone supplies it | Not scrapable; needs a person |
| Whether the phase-inversion claim generalises | Contra's diagnostic pitch rests on it, on 28 evaluators and 80 sessions with no multiple-comparison correction | Replication; also a publishable result |
| Whether app-layer companies buy through unposted channels | Zero vendor mentions across 1,562 postings could mean no purchasing, or purchasing that never touches a job description | Ask Canva directly |
| Whether a final USCO Part 3 exists | The pre-publication text is the only version, from an office whose Register was removed days later | Docket check |
| Whether the Digital Omnibus left Chapter V untouched | The Article 53 selling argument rests on it | Read Regulation (EU) 2026/1744 in full |
| Nordic moral-rights waivability outside Sweden | Denmark, Norway and Finland are assumed parallel and were not verified | Local counsel, one memo |
| A benchmark for portfolio-review cost at scale | Nobody has published what the human screening layer costs | Measure your own and publish it |
No copyright, database-right or trade-secret case anywhere has been decided on a commissioned expert-judgement corpus. Every legal conclusion in this dossier applies adjacent doctrine to a fact pattern no court has seen. That is not a research failure — it is the honest state of the field, and it is why the constraints pages point at primary material for counsel rather than at conclusions.
Two research habits are worth carrying out of this list. First, the most valuable unknowns here are cheap — a survey, a re-analysis, a handful of emails — and they are unknown because nobody has bothered, not because they are hard. Second, three of the seven top items close through conversations rather than searches, which is the strongest argument for treating the people map as a research instrument rather than as a sales list.