The single most useful correction to the "publish a design benchmark and own the category" plan is that somebody already did it, and it was not a design company.
UI-Bench is the most rigorous public human-judged design benchmark in existence, and it was built by AfterQuery — an expert-data company — roughly a year before Contra Labs launched. The paper is authored by Sam Jung, Agustin Garcinuno and Spencer Mateega of AfterQuery, with UPenn affiliation (arXiv 2508.20410).
What it measures, and who it measures
UI-Bench benchmarks exactly the buyer segment a design-data vendor targets: "v0 by Vercel, Bolt, Orchids, Lovable, Replit, Figma Make, Base44 by Wix, Magic Patterns, Same.new, and Create Anything" (arXiv 2508.20410).
| Element | UI-Bench |
|---|---|
| Method | Blinded expert pairwise comparison, TrueSkill with confidence intervals |
| Scale | 30 prompts, 300 sites, 4,047 pairwise matches |
| Panel | 194 expert judges — designers 65.5% (127), web developers 57.2% (111), researchers 60.3% (117), other 7.2% (14); overlapping categories |
| Prior exposure | only 10.8% had used AI website-generation tools before |
| Result spread | Orchids 30.12 (67.5% win rate) → Replit 20.95 (38.9%) |
| Openness | prompts, evaluation framework and leaderboard published (leaderboard) |
TrueSkill over Bradley–Terry is a deliberate methodological choice: it yields calibrated confidence intervals rather than point estimates, which is what lets the paper say two tools are not distinguishable rather than pretending to a ranking.
The benchmark exists because the thing it measures is real: these tools share underlying models and still produce visibly different design quality, and until UI-Bench nobody could say by how much. That spread — a ten-point TrueSkill gap between the top and bottom of the app layer — is the closest thing to a demand signal in the app-layer buyer map, and it is worth noting that the tool sitting at the bottom of it, Replit, is run by a CEO who is himself a named angel investor in design-AI research — Amjad Masad appears on Lica World's investor list (lica.world).
The panel appears to have been free
Recruitment, verbatim: "We hand selected and invited the experts in this study on the basis of having professional UI and UX experience as a designer, web developer, researcher, or other relevant profession." And the acknowledgement: "We thank the expert panel for generously contributing their time and judgment; their participation made large-scale, blinded data collection feasible" (arXiv 2508.20410).
There is no compensation statement anywhere in the paper. "Generously contributing their time" is the standard formula for an unpaid panel, and that is how it reads — but the paper does not say so, and this is an inference rather than a disclosure. [WEAK] on "unpaid"; [HARD] on "not disclosed".
If the inference is right, AfterQuery assembled 194 design professionals at zero cash cost and produced the benchmark that defines design quality for the app layer. Compare the alternatives on the same axis: HEIM's crowd panel cost $13,433.55 at $16/hr (arXiv 2311.04287); GenAI-Bench's about $9,600 at $12/hr (arXiv 2406.13743); Contra's TASTE study about $13,000 at $90/hr (arXiv 2605.20731). The most expert panel in the set was the cheapest.
AfterQuery is not a small player using a stunt. Its own site states it has "closed $30M Series A at $300M valuation and surpasses $100M revenue run rate," positions as "Expert Data for Frontier AI," and claims to be "Powering every frontier AI research lab" (afterquery.com) — all self-reported, and none of it broken out by vertical.
The lesson: benchmarks here are customer acquisition, not product
The play is legible because two companies have now run it. An expert-data vendor publishes an open, rigorous, human-judged benchmark in a vertical, establishes methodological authority, and sells the private version to the companies the benchmark ranked. Contra Labs ran it with the Human Creativity Benchmark and 39 studies in five months; AfterQuery ran it in the design vertical about a year earlier, from a generalist position, with a bigger panel and better statistics.
"Publish a benchmark to establish credibility" is not a differentiator in this category. It is the entry ticket, it was played first by a generalist, and any design benchmark the client builds should be booked as a marketing cost with a citation return, not as a revenue line. Publishing the benchmark is the general case; Where the headroom is is the design-specific one.
Two second-order consequences follow, and both are uncomfortable.
A published benchmark is a spent asset. Its defensibility decays the moment it is released, because the method is the product and the method is now public. The only resolution is a two-tier corpus: a public tranche sized to generate citations and inbound, and a permanently sealed tranche that is never released and is the thing actually sold. Eight units and one cost anchor carries the sellable-unit design.
The panel cost is the wrong thing to optimise. UI-Bench demonstrates that senior design professionals will contribute judgement for free when the ask is small, blinded, credited by association and intellectually interesting. That undercuts any pitch built on "we can recruit designers" — so can a three-author paper. It also sharpens where the money actually goes: recurring, high-volume, unglamorous evaluation work is what needs paying for, and Will they say yes is where the recruiting argument actually lives.
The counter-move is cheap and nobody has taken it
When the benchmark's author is a commercial party in the same market, the benchmark is not neutral infrastructure. Buyers increasingly know this. Note the structure carefully and without allegation: the top-ranked tool in UI-Bench, Orchids, holds the largest margin over the field. There is no evidence of any relationship between AfterQuery and Orchids and none is alleged here. The general point stands regardless of this specific case — a vendor-authored leaderboard carries an unpriced conflict, and every vendor in this category now has one.
That leaves an unoccupied position, and it costs almost nothing to take:
- An independent methodological advisory board — named academics and practitioners with published terms of reference, who can veto a methodology change.
- Conflict-of-interest disclosure as standing practice — every commercial relationship with any ranked party, published beside the leaderboard, including "none".
- Pre-registration of evaluation protocols — the prompt set, the rubric, the rater-selection rule and the analysis plan published before results are collected, so a benchmark cannot be retrofitted to a flattering conclusion.
None of this requires capital. All three are governance, not engineering. And they are exactly the properties an institutional or European buyer responds to, which is the argument The one right a US competitor cannot hold makes in full and Sell the paperwork with the data turns into a sellable artefact.
AfterQuery already owns the credibility asset in design benchmarking, from a $100M+ generalist base, using a free panel. Do not try to out-publish it. Out-govern it: the first vendor whose benchmark carries pre-registration and a real COI disclosure owns the only differentiator left in a market where everyone publishes.
AfterQuery's panel compensation, whether UI-Bench generated any commercial relationship with the ranked vendors, and whether any of its $100M+ run rate is design work rather than code and finance. All undisclosed; the revenue and valuation figures are self-reported on the company's own site with no independent corroboration.