This is the page that decides whether this is a business or a consultancy, so it is worth stating the problem in its harshest form before answering it.
In offensive security the artefact has a perfect oracle: the exploit fires or it does not. A perfect oracle makes a product — you sell a graded, verified capability set and the buyer checks it without you. No oracle makes a permanent human service — you sell the judgement of a person, repeatedly, forever, and your revenue is your headcount times your rate. That is the difference between a data business and a staffing business, and everyone in expert data sits somewhere on that line.
Aesthetic judgement has no oracle at all. Not a noisy one, not an expensive one, not a slow one. There is no procedure, at any price, that returns whether this logo is better than that one. The naive conclusion is that a taste-data company is a consultancy forever.
That conclusion is wrong, and the reason it is wrong is the most important argument on this site.
The manufactured oracle
Where nature supplies no ground truth, an institution can manufacture one. This is not a novel move — several large industries already run on it.
- Credit ratings. There is no oracle for "will this bond default." Moody's and S&P manufacture one by publishing a methodology, applying a committee, and being consistent for decades.
- Clinical reference standards. Where no gold standard exists, medicine uses adjudicated expert panels and reports the agreement statistic alongside the label (Centaur.ai, in full is the commercial version of this).
- Wine, gymnastics, architecture competitions. Panel plus rubric plus published protocol.
The manufactured oracle has exactly three components, and all three are ownable:
- A panel — a named, vetted, stable set of judges. Stability matters more than size past a point, because a rotating panel makes longitudinal scores incomparable, which destroys the only thing a buyer wants from a repeat purchase.
- A protocol — the rubric, the anchors, the adjudication rule for disagreement, the exclusion rules, and the version history of all of them.
- A published reliability figure — the number that says how much to trust the label.
A buyer cannot rebuild any of the three from a delivered file. They can rebuild the labels — labels are just values in a column. They cannot rebuild the panel's identity, the protocol's calibration history, or the reliability track record. That is the moat, and it exists because there is no oracle rather than in spite of it. In a domain with a real verifier, the panel is redundant and the file is the whole product.
Say it the other way round, because this is the sentence that carries the business: a single opinion has no error structure. A calibrated panel does. One art director's verdict is an anecdote and the buyer knows it. Nine art directors with a measured Krippendorff's α, a documented exclusion rule and a trimmed mean is an instrument — it comes with a confidence interval, and a confidence interval is a thing an ML team can put in a loss function.
Measured disagreement is the artefact, not the noise
Contra's central intellectual claim is that evaluator disagreement in creative domains is signal, not noise — that it distinguishes convergence (shared professional standards: typography, layout, hierarchy) from divergence (legitimate taste variation) (arXiv 2606.30561). This is correct, and it is under-exploited, including by them.
The academic foundation predates them by years. Vessel et al. (2018), Cognition 179, decompose aesthetic response into shared and individual components and show the split is domain-dependent — high shared taste for faces and landscapes, high individual variation for architecture and artworks (Max Planck Neuroscience). Design sits on the "artifact of human culture" side of that divide, which predicts high individual variation — and that is exactly what Design Crit's 74.1% human ceiling shows empirically (contralabs.com/research/design-crit).
Three products fall out of taking disagreement seriously, and none of them is a label:
The convergent core. The subset of judgements on which professionals agree at high reliability. This is the safe training signal — the part a lab can optimise against without teaching the model one person's taste. It is smaller than the full corpus and worth more per unit.
The divergence map. Where professionals legitimately split, and along what lines: discipline, geography, seniority, training. This is what lets a lab build a steerable model rather than one with a single averaged taste. No competitor sells this as a distinct SKU and it is the most differentiated thing in the taxonomy.
The reliability certificate. Per axis and per phase, the achieved agreement coefficient and the panel depth used. It makes the other two auditable, and it maps directly onto AI Act Article 10-style documentation obligations (The one right a US competitor cannot hold).
Note what this does to the 74.1% ceiling. Read naively it is a damning number: professionals agree only three times in four, so what exactly are you selling? Read correctly it is the product spec. The 74.1% is not the failure rate of the panel; it is the measurement that makes the panel usable, because a buyer who knows the human ceiling can tell the difference between a model that is genuinely wrong and a model that has landed inside the zone where experts disagree. Nobody can do that without your reliability figure. See Where the headroom is for the numbers on both sides of that gap.
The gap that sets the entire cost structure
Everything above is theory until you answer one operational question, and the field has not answered it.
No study establishes, for professional design evaluation specifically, the panel size needed to reach a stated reliability target per axis. What exists is practice with no published basis: Contra uses "3+ professional creative evaluators" per output (contralabs.com/human-creativity-benchmark); VAB uses 10 independent expert judges per task (arXiv 2605.12684); Design Crit uses two cohorts of five; Awwwards uses a minimum of 18 jurors per site with the three most extreme scores discarded (Awwwards). That is a six-fold range with no stated derivation anywhere.
The two most directly relevant adjacent sources — the Konstanz expertise-screening work on crowd image-quality assessment, and the CrowdMOS methodology — were not retrievable this session (robots-blocked), so even the imported answer is unavailable. [UNVERIFIED]
This is the most valuable gap in the dossier and the cheapest to close. It is a re-analysis of your own pilot data, not new collection. It sets your entire cost structure — panel depth is the single largest driver of unit cost in every artefact you sell. It is publishable. And no competitor has published it.
The theory you need is not exotic. Reliability of a panel mean follows the Spearman-Brown relationship: to reach a target reliability you need roughly n raters, where n scales inversely with single-rater reliability. Since single-rater reliability is axis-dependent — HCB reports agreement highest on prompt adherence and lowest on visual appeal — and phase-dependent — Kendall's W of 0.345 at ideation rising to 0.549 at refinement (arXiv 2606.30561) — the required n varies across a two-dimensional grid that nobody has filled in.
Filling it in changes the business, not just the paper. It converts panel depth from a guess into a dial, lets you quote different prices for different confidence levels, and gives a sales conversation an answer to "why nine raters and not three" that is a number rather than an assertion. It is the same move Phase decomposition argues for on the pricing side, and it is the first thing to run in the first ninety days.
So: business or consultancy?
It is a business, on three conditions. Fail any one and you are an agency with good margins for a while.
One — the panel must be a stable, named asset, not a labour pool. A rotating crowd cannot manufacture an oracle, because the thing being sold is consistency over time. This has a direct consequence for the contributor agreement: you need continuity, which argues for retainers or panel membership rather than pure per-brief piecework. Awkwardly, retainers push toward employment-classification risk in the EU, so the two constraints trade off against each other and someone has to own that trade.
Two — the protocol must be versioned, documented and treated as a trade secret. It is the component with the longest build time and the highest replication cost, and unlike the labels it is authored by you rather than licensed from a contributor (Copyright is the weakest thing you own).
Three — delivery must be a subscription against a sealed suite, not a file transfer. A file transfer is a consultancy invoice. A sealed suite with quarterly refresh is a product, and it is the only unit where the missing oracle is an advantage: because ground truth is your panel's adjudicated judgement and your panel is private, the buyer cannot reproduce your scores without you.
Contra currently fails condition three. Its job posts describe owning "scope, staffing, timelines, and client relationships" and converting "unclear client questions into well-defined research questions" (Ashby). That is agency language, and it is the gap to attack — the argument developed in What to build first and scored in The taste read.
One closing comparison, because it disciplines the optimism. In security, the logic-bug wedge exists precisely where the oracle fails — business-logic flaws that no scanner can verify — and that is the most durable human work in the category. In health, adjudicated panels with published agreement statistics are how the entire diagnostic field handles the absence of a gold standard. The taste market is not unusual in lacking an oracle. It is unusual in that nobody has yet built the instrument that measures the absence.