The company to build is not a taste-data company. It is a calibrated, named, consented professional panel, and the held-out evaluation content that comes out of it. The panel is the asset. The files are inventory the panel produces, and they should be priced like inventory.
That sounds like a distinction without a difference until you ask what a buyer can reconstruct from what you deliver. Labels are values in a column and any competent team can reproduce them. What they cannot reproduce is the panel's identity, the protocol's calibration history and the reliability record — which means the only durable thing you can sell is the instrument, not its output. The taste read argues why every other shape in this lane is already occupied or already free. This page is the shape that is left.
Agreement is the manufactured oracle
In offensive security the artefact has a perfect oracle: the exploit fires or it does not, and a perfect oracle makes a product because the buyer can check it without you (The one shape that survives). Aesthetic judgement has no oracle at all — no procedure at any price returns whether this logo is better than that one. The naive conclusion is that this is a consultancy forever.
Where nature supplies no ground truth, an institution manufactures one. Credit ratings have no oracle for "will this bond default"; Moody's manufactures one by publishing a methodology, applying a committee and being consistent for decades. Medicine does the same where no gold standard exists, using adjudicated panels and reporting the agreement statistic beside the label (Characterised disagreement). The manufactured oracle has three components — a stable named panel, a versioned protocol, and a published reliability figure — and all three are ownable. The oracle problem works through it.
A single opinion has no error structure. A calibrated panel does. One art director's verdict is an anecdote and the buyer knows it. Nine art directors with a measured Krippendorff's α, a documented exclusion rule and a trimmed mean is an instrument — it arrives with a confidence interval, and a confidence interval is something an ML team can put in a loss function.
This is also why measured disagreement is the artefact rather than the noise. Design Crit puts the human ceiling at 74.1% — two professional designers agree on which of two images is better about three times in four (Design Crit). That is not a failure of the panel; it is the map. It separates the convergent core where professionals agree at high reliability, which is the safe training signal, from the divergence map showing where they legitimately split and along what lines — discipline, geography, seniority. Nobody sells the divergence map as a distinct SKU, and it is the most differentiated unit in the taxonomy. It is also the only thing that lets a lab build a steerable model rather than one with a single averaged taste, which is exactly the failure Taste Labs' own manifesto names when it says reward models learn "what nobody dislikes, rather than what anyone loves".
Why the panel holds against the arenas
Arena's image board carries 6,047,075 votes collected free (arena.ai). You will never approach that and should never try. But the arena's four structural absences are permanent, because they are properties of the product form rather than of its scale.
No identity. An arena vote is anonymous by construction; the demographic cut Design Arena sells is a continent, not a credential. A panel can say this judgement came from an ADC Product Design juror with fourteen years in brand identity, and can say it per record.
No rationale. A click cannot carry "the baseline grid breaks at the third module" or "the optical sizing is wrong at display weights." That vocabulary is what makes a critique usable for supervised fine-tuning and for training a critic rather than a scorer — and MJ-Bench found that natural-language feedback from judges is more accurate and stable than numerical scoring, which is a published result pointing straight at the written unit (arXiv 2407.04842).
No axes. An arena collapses everything into one winner. HPSv3 has two dimensions; Design Crit has nine. The specialist play is not to out-scale the crowd datasets but to occupy the axes they collapse, because a lab that knows only which output lost has learned nothing it can act on.
No process. No arena records how the work was made. That is the trajectory asset, and as of August 2026 the narrated-professional-creative-trajectory category has one public occupant whose largest dataset is four trajectories.
Why it holds against the generalists
The generalists have the supply and lack the discipline. Mercor pays $150–250/hr for senior design reasoning capture and $80–150/hr for agency-pedigree rubric authoring and scored critique (Mercor) — but it runs the whole creative vertical with generalist project leads and no design staff on its own board, and it publishes no reliability statistic for any of it. Handshake distributed $100 million to 750,000+ fellows and has no design vertical at all (Handshake AI). Contra uses "3+ professional creative evaluators" per output; the Visual Aesthetic Benchmark used ten independent expert judges per task; Awwwards runs a minimum of eighteen jurors with automatic rejection of the three most-outlying scores (Awwwards). That is a factor-of-six spread in panel depth across the field with no published basis for any of it.
Nobody in this market publishes panel composition, recruitment channel, pay, per-axis reliability, or the exclusion rule. The AI Act is simultaneously making documented provenance a compliance artefact. A vendor that publishes all six converts a cost into a differentiator, and it is the one differentiator a generalist cannot copy without rebuilding its operating model.
The agreement gradient is the finding you can own
Contra's phase decomposition has two halves and the field is quoting the wrong one. The eye-catching half — that model rankings invert between ideation, mockup and refinement — rests on 28 evaluators across 80 sessions with no correction for multiple comparisons, and inversions expire with each model release. The half that generalises is this:
On HCB ad images, Kendall's W runs 0.345 at ideation → 0.436 at mockup → 0.549 at refinement (arXiv 2606.30561). The authors' own gloss: criteria become "progressively more verifiable" as work advances. It has never been replicated.
That is a structural claim about creative work rather than about any model, so it does not expire — and it is the only claim in this dossier that is simultaneously a methodological contribution, a cost advantage and a sales argument. Because if reliability is a function of phase, then cost per unit of reliable signal is a function of phase, and panel depth should vary accordingly: deep panels early where agreement is low, thin panels late where it is high. The same logic runs across axes, since agreement is highest on prompt adherence and lowest on visual appeal.
That is a pricing schedule that falls out of a finding, and it makes you an instrument maker rather than a labour broker. Phase decomposition has the mechanics. The replication is available to whoever runs it first, and it is a re-analysis of your own pilot data rather than a new collection — which is why it sits at number two in What the taste dossier could not establish.
De-novo briefs are the only structurally safe substrate
Contra records trajectories during real client briefs and treats that as the differentiator. It is a liability. A screen recording of a real project captures the client's brief, unreleased brand assets, an unlaunched campaign and possibly their customer data, in frame. The contributor cannot license what they do not own, and the injured party is a client with no contract with you, no notice, and a straightforward claim in confidence and in copyright.
De-novo briefs — written for the purpose, never web-published — cost realism and buy three things. They are the only cleanly sellable substrate for trajectory capture at scale. They are the only material that is provably uncontaminated, which matters increasingly to a buyer who must assume every public benchmark leaked into pre-training. And they are the natural substrate for the sealed suite, which is the highest-margin unit in the taxonomy because it converts a data sale — one-time, resellable by the buyer, margin-compressing — into a subscription whose switching cost grows with the length of the buyer's own score history.
The complementary control is in The clause that expires your corpus in year ten: an employer-and-client warranty with an onboarding attestation behind it, a positive-consent step per submission, a pre-submission redaction gate, and exclusivity on the artefact but never on the person.
Provenance ships with the data
Article 53(1)(d) of the AI Act requires GPAI providers to publish a sufficiently detailed summary of training content on the AI Office's template, and 53(1)(c) requires a copyright policy identifying rights reservations "including through state-of-the-art technologies" (artificialintelligenceact.eu). Enforcement powers went live on 2 August 2026 with fines to 3% of global turnover, and compliance is visibly patchy — Anthropic, Mistral and xAI substituted narrative prose for the template.
Nobody is failing this because the boxes are hard to fill. They are failing because web-scraped data of unclear provenance is the hardest thing to describe honestly and the most embarrassing thing to publish. A commissioned, consented, contract-backed corpus is the easiest line a lab will ever write into that summary: named channel, dated agreements, per-record provenance, a consent basis, a jurisdiction. Bundle it — provenance file, consent records, licence chain, template-ready language — and price it in. It costs almost nothing because the database right already requires you to keep the ledger. Sell the paperwork with the data has the full argument, and Copyright is the weakest thing you own explains why Bartz priced acquisition rather than use, which is what makes this a risk transfer rather than a quality claim.
What to build first, concretely
Run it in product and UI design, not in image aesthetics. Four reasons, in order. The buyer's user is a professional, so the Pick-a-Pic objection — that real users are better ground truth than experts — does not bite, because professional acceptability is not consumer preference. The gap is measured in public by a third party: UI-Bench spreads Orchids at 30.12 (67.5% win rate) down to Replit at 20.95 (38.9%). The buyers are named, numerous and ranked against each other, with $31.6B of vibe-coding valuation employing nobody to evaluate design quality (The creative tool layer). And the substrate authors itself: a landing-page brief is not confidential, where a brand identity brief always is.
The artefact. Forty de-novo briefs, each executed by four systems, evaluated in three phases — wireframe, comp, build — across five axes, by a panel of about twenty-four raters seeded from the Awwwards, ADC and D&AD juror rosters and ADPList, screened attitudinally before any skills test and then at roughly 69% agreement with expert consensus on a held-out ranking set, which is the human-expert score on the Visual Aesthetic Benchmark and the only published pass mark in creative supply (Telling a good designer from a confident one). Collect rankings, not scores — VAB found direct ranking yields substantially higher inter-annotator agreement than rankings derived from scores, and it is free to adopt. Panel depth varies by phase: nine at ideation, five at mockup, three at refinement. Ship the data plus a reliability dossier — panel register with credentials, protocol version, per-axis and per-phase agreement coefficients, exclusion log, adjudication trail. Then build an identically constructed mirror set and never release it.
What it costs. Forty briefs at the $150–600 band is $6K–24K of senior authoring. Calibration is 24 raters × ~3 hours at $75–100/hr, so $5K–7K. Judgement production of roughly 5,000 ranked comparisons at $0.91 each plus written rationale on one in five at $8–25 each is about $17K. Verification — adjudication, exclusion, reliability computation — adds 15–20%. Call it $35,000–55,000 of direct expert labour for the first suite, and $70,000–110,000 for the pair including the sealed mirror [INFERENCE — the time and depth assumptions are mine; the rates and the $0.91 anchor are observed]. No rigs, no licences, no data-use agreements, no security gate before the pilot.
What it prices at. The nearest public anchor is Mercor's APEX: over $500,000 for 200 expert-authored benchmark tasks, about $2,500 a task (Time). Forty briefs across three phases is 120 task instances. The only other bound is structural: a bespoke Contra Labs engagement plausibly prices at $75K–300K on a cost-plus-margin reconstruction. Plan on $150,000–300,000 for the first suite and $200,000–400,000 a year for the sealed subscription with quarterly refresh — and state plainly that no rate card exists anywhere in this market, from any vendor, for any artefact, which is both the risk and the reason the first credible one sets the price.
That is a 65–80% gross margin on direct labour. What compresses it is research-lead time per engagement — Contra's flat $150–200K research band is a consultancy cost structure — which is the whole argument for a suite built once and sold repeatedly rather than a bespoke study each time.
The trap
The corpus business. It looks like the same business and it is not. A frontier-grade 1.17M-comparison preference corpus is roughly $1M of direct annotator labour at the $0.91 rate the competitor published about itself. That is a line item a lab approves in one meeting, and several already have. Rows have a going rate, the rate is set by whoever has the cheapest credentialed labour, and it falls every year. The durable thing is the panel and the calibration, not the file.
Two smaller traps sit behind it. A buyer who wants the panel's answer at a single-rater price, framed as simplification — saying no to that is the strategy, because a single reader produces a label and a panel produces an instrument, and only one of them is repeat business. And publishing everything: Contra gave away eight datasets under CC BY and taught the method for free, while AfterQuery already owns the benchmark play in this exact niche. Publish the panel-size curve, which is a method nobody has. Keep the sealed suite sealed.
The sequence for testing all of it, with a disproof attached to every step, is at Ninety days in taste; the assumptions it rests on are at What the taste dossier could not establish; the people who can close them are at Seventy people and one missing bridge and [/people]. The same argument in the two other lanes is at The one shape that survives and Characterised disagreement, under The security read and The health read, with the general form at The specialist wedge and Expert data for frontier labs.