Build one thing first, and make it this: a proper rebuild of the Human Creativity Benchmark whose published contribution is the reliability methodology, not the leaderboard.
That sentence contains a choice at every word, and the choice is arguable at every word. What follows is the argument, then the design, then the cost, then the things that could make it a mistake.
Why rebuild something rather than invent something
The steer for this whole section is prefer taking something that already exists and doing it properly over inventing a new category. That is not modesty. It is what the scoreboard says works.
The evidence is that a new category bought nobody anything. Contra invented "the Human Creativity Benchmark" and "Creative Arena" as named categories and holds zero citations on both. Lica invented a benchmark taxonomy across four papers and holds ten citations of which eight are its own. What travelled instead was FinanceQA — 148 rows, an obvious format, a repo, 17 citations — and UI-Bench, which is not a new category at all but a rigorous version of a thing that already existed. A new name has to earn attention before its content can be judged. A better version of a known instrument is judged immediately, against something the reader already understands.
So the question is only which existing thing.
Why HCB is the target
Four reasons, in order of weight.
It is the flagship benchmark of the company being modelled. The whole dossier is built around whether Contra Labs' shape can be run better in an adjacent position. HCB is the artefact that carries their claim to methodological authority — the arXiv paper, the phase decomposition, the disagreement-is-signal argument, the "industry standard for measuring creative AI" copy. Doing it properly is the most direct possible statement of the alternative.
It is completely public, which makes the re-analysis cheap. contralabs/HumanCreativityBenchmark is CC-BY-4.0 and holds 8,012 rows across five CSVs: 95 prompts, 380 model outputs, 3,174 pairwise comparisons, 2,116 scalar ratings, 2,247 qualitative responses. Prompts, outputs and judgments are all there. Nothing is held out. A re-analysis needs no permission, no partnership, and no new collection to get started.
It is small enough to be beaten and serious enough to be worth beating. The paper reports 28 evaluators from 13 countries, 93 prompts, 80 sessions, 13 models across 5 domains × 3 phases, with Bradley–Terry Elo, Kendall's W, Krippendorff's α and Friedman tests. That is real methodology at a scale two people and a budget can exceed. Their own dataset card says so: "The evaluator pool is modest (31 designers…) and prompts were sampled once. Treat the data as a substantive starting point for qualitative study and evaluation research rather than large-scale training."
And it has zero citations. The ground is not contested. Nobody has replicated it, extended it, or argued with it in public.
The paper describes 5,940 pairwise, 5,940 scalar and 3,675 qualitative judgments — 15,555, rounded in the abstract to "15,000 professional judgments." The public dataset contains 3,174 pairwise, 2,116 scalar and 2,247 qualitative = 7,537, and the card says 31 evaluators and 95 prompts where the paper says 28 and 93. That is 48% of the described judgments, with an evaluator count that moves in the opposite direction from what dropping incomplete responses would predict. Neither document explains it. Any technically literate reader finds this in ten minutes — which means you must state it once, neutrally, in a footnote-shaped sentence, and never build a rhetorical beat on it. See the risks section.
The specific defect to attack
Not the sample size. Sample size is a criticism anyone can make and nobody is paid for.
Nobody has published how many designers per item you need before an aesthetic ranking stabilises, and how that number varies by criterion. Not Contra, not Lica, not AfterQuery, not Design Arena, not the academic aesthetics literature. Practice ranges from Contra's "3+ professional creative evaluators" to the Visual Aesthetic Benchmark's 10 expert judges per task to Awwwards' minimum of 18 jurors — a six-fold spread with no published derivation anywhere.
Two pieces of external evidence make this the right target rather than merely an open one.
Google says the common practice is wrong. Forest vs Tree: The (N,K) Trade-off in Reproducible ML Evaluation (Flip Korn and Chris Welty, AAAI; blogged 31 March 2026, with an open simulator at google-research/vet) finds that "the common practice of using 1, 3 or 5 raters per item is often insufficient" and that practitioners "often need more than 10 raters per item." That is directly adverse to Contra's published "3+ evaluators" rubric and to the 8–12-per-study panels in their own methodology page. It is also domain-general: Korn and Welty did not run it on aesthetic judgment, which is the highest-variance case there is.
The right statistical machinery already exists in this market and was pointed at the wrong question. TASTE (2605.20731) validates its rankings with per-prompt Kendall's τ, majority-vote probability pmax and Condorcet cycle rate against exact iid-uniform nulls — precisely the apparatus a panel-size sweep needs — and then runs the entire study at a fixed R = 5 and never sweeps panel size. Two disjoint cohorts of five designers, nine criteria, 80 prompts each. The tool is built. It was used once, at one depth.
So the artefact is: the panel-size-to-reliability curve for professional design evaluation, per criterion. It is a re-analysis and a targeted collection, not a new field. It answers a question a named competitor's published rubric gets wrong, using a method a second competitor built and did not apply. And it produces a number every buyer in this market needs and none can currently get.
The design, concretely
Sweep panel size. Run R = 1, 3, 5, 7, 9, 11, 15 raters per item, on the same items, and report for each criterion the reliability curve and the point at which it stabilises — the depth beyond which adding raters no longer moves the aggregate ranking by more than a stated tolerance. Report it as a curve with confidence bands, not a single recommended number. The comparability requirement is doing this on the same items across all depths, so the criteria can be read against each other rather than against seven different studies.
Report at minimum: Krippendorff's α, Kendall's W, pmax, Condorcet cycle rate, and rank stability under resampling at each depth. TASTE's own α of ≈0.19 on rankings — against DesignPref's 0.25 with 20 designers — is the honest baseline this field starts from, and a curve is far more useful than another single α.
Use rankings, not scores. The Visual Aesthetic Benchmark's central methodological finding is that direct ranking by experts yields substantially higher inter-annotator agreement than rankings derived from individual scores (arXiv 2605.12684). It is free to adopt and it improves every number downstream. HCB collects scalar 1–5 ratings alongside forced-choice pairwise; the rebuild should lead with ranked comparisons and treat scalars as secondary.
Screen the raters against a published pass mark. VAB puts human experts at 68.9% on identifying both best and worst image correctly, against 26.5% for the best frontier system. Screen applicants at roughly 69% agreement with expert consensus on a held-out ranking set. That is the only externally-sourced, published pass mark available in creative supply, and using it means the panel's composition is defensible in a sentence rather than an appendix (Telling a good designer from a confident one).
Keep a sealed test split behind a submission API, and publish the dev set. Every design benchmark in this market is fully public and therefore contaminated: HCB's prompts, outputs and judgments are all in the CC-BY repo; UI-Bench's prompts are on Hugging Face; GDB is a 17 GB public download; Design Arena's prompts are user-generated and never published at all. Nobody operates a held-out creative test set with a submission API — the standard structure in every mature benchmark category. Publish enough to be replicated and cited; seal enough to still be measuring something in a year.
Fix the instrument so it can be re-run. Contra's 39 studies use different panels, different prompts and different rubrics each time, so they cannot be stacked into a trend line. Nobody in this market re-runs the same instrument against successive model releases and publishes the series. A frozen protocol, a frozen prompt set, a stable panel and a versioned rubric turn a one-shot snapshot into a longitudinal instrument — which is also the only version of this that is a subscription rather than a paper.
Pair it with open scorer weights. Zero design-domain reward-model checkpoints are public from any organisation in the inventory. Lica's taste-scorer ships MIT-licensed inference code and the data, on a frozen Qwen3-VL-Embedding-2B backbone reaching 0.611 pairwise accuracy against a 0.741 single-rater ceiling — and no trained checkpoint anywhere on Hugging Face under purvanshi, lica-world or contralabs. Publishing weights costs one upload and is the single cheapest way to be the artefact people build on rather than the artefact people cite once.
What it costs
The anchors are disclosed, which is unusual and worth using.
The Contra × Lica TASTE paper discloses ~$90 per hour, 10 designers in two cohorts of five, ~13 and ~16 hours each (≈14.5 average), for ≈$13,000 of direct annotator labour. HEIM — Stanford CRFM's multi-axis, aesthetics-inclusive text-to-image evaluation — states its total human-annotation bill outright at $13,433.55. Two field-defining efforts, both around thirteen thousand dollars.
$13,000 against the 14,460 ranking rows published in the HF repo is ≈$0.90 per row — the $0.91 anchor used throughout Eight units and one cost anchor and The taste read. $13,000 against the 21,600 underlying pairwise judgments the paper's design implies (80 prompts × 5 raters × C(4,2) = 6 × 9 criteria) is ≈$0.60. Both are defensible; they are not the same number. Quote the denominator every time.
A swept version costs more by construction, because the whole point is more raters per item. Rough shape, stated as planning bands rather than quotes:
| Line | Basis | Band |
|---|---|---|
| Rater time for the sweep | ~50 screened raters × 12–20 hrs × $90 | $54K–$90K |
| Screening and calibration | ~120 applicants × 1–2 paid hrs | $9K–$20K |
| De-novo sealed split authoring | 40 briefs at $150–600 | $6K–$24K |
| Analysis, harness, write-up | 1 person, ~10 weeks | in-house |
| Direct expert labour, total | $70K–$135K [INFERENCE — the depth and hour assumptions are mine; the $90/hr and the HEIM total are observed] |
That is five to ten times what TASTE and HEIM each cost, and it is the right multiple, because the deliverable is the thing neither of them produced: a curve across depths instead of a point at one depth. What it buys is a number that sets your entire cost structure — panel depth is the largest single driver of unit cost in every artefact in the taxonomy — plus a pricing argument that is a measurement rather than an assertion, plus the only methodological position in this market that a generalist cannot copy without rebuilding its operating model.
What to ship
Six things, together, on one day.
- The dataset, CC-BY-4.0, with the dev split, the full judgment records including rater ids and depth labels, and a datasheet naming panel composition, recruitment channel, pay and screening threshold.
- The paper, on arXiv, leading with the reliability curve and not with a model ranking.
- A GitHub repository with the harness, the scoring code, the reliability computations and the resampling procedure, under a permissive licence. This is the non-negotiable one.
- Open scorer weights on Hugging Face, with the training script.
- A leaderboard that is a live page with dates on it, not a table in a PDF.
- The sealed split, described publicly, held privately, reachable through a submission endpoint.
Every cited AfterQuery artefact has a repository. Every uncited Contra artefact does not. Contra has no GitHub organisation at all — for a company whose product is evaluation, that means there is no way for a lab to run a Contra evaluation. Shipping the harness is not a nice-to-have on this release; on the evidence in What publishing actually bought them it is the variable.
The risks, honestly
It is a re-analysis of a competitor's work and must be scrupulously fair to them. The defect being attacked is a field-wide practice that Contra happens to have written down more explicitly than anyone else; that is a reason to credit them, not to hit them. Cite HCB as the reason the re-analysis is possible at all — they published the data, which almost nobody in this market does. State the 48%-of-judgments discrepancy once, in neutral language, as a data-availability note. Do not lead the abstract with it, do not put it in a title, and do not tweet it. A paper that reads as an attack on a named competitor will be discounted by exactly the audience it is written for, and it forecloses the collaboration that is the better long-run outcome.
The result might come out against the client's own pricing. If the curves stabilise at 11 or 15 raters per item on the divergent criteria, then honest panels are expensive, and the number the client publishes becomes the number a buyer uses to argue that the client's own quotes are too thin. That risk is real and should be accepted anyway: the alternative is competing on an unpublished assertion about depth, which is the position everyone already occupies and nobody wins from. If the curves come out low, the finding is an even better sales asset. Either way, publishing the instrument is compatible with charging for the panel — The oracle problem is the argument for why the method can be public while the panel is not.
And a benchmark is customer acquisition, not product. This is the most important caveat on the page. The artefact gets meetings, citations and a defensible methodological position; it does not, on its own, produce revenue, and the record shows no company in this cohort tying a publication to a disclosed deal. The business it is meant to open is the calibrated panel and the sealed suite argued at What to build first, with the oracle argument at The oracle problem and the phase-and-axis pricing schedule at Phase decomposition. Build this first because it is the cheapest credible entry into that conversation — not because it is the business.
The ranked list of everything else that does not exist is at What nobody has built; how to make this release actually travel is at How an artefact travels; what could not be established about the build stack is at What this plan could not establish; and the ninety-day sequence it fits inside is Ninety days in taste.