Miju Labs

The security dossier

Where the headroom is

Enormous headroom in taste, none in compositionality: the best frontier model scores 26.5% on the Visual Aesthetic Benchmark against 68.9% for human experts, while GenEval drifted to 17.7% absolute error and had to be replaced. Half the design benchmarks that matter were built by companies selling into the same market.

high confidence9 minupdated 2026-08-30benchmarks · evaluation · saturation · vendor capture · ui-bench · vbench

Benchmarks matter here for two reasons that pull in opposite directions. They tell you where model capability is weak, which is where a buyer's budget goes. And they are the field's dominant marketing instrument, which means several of the most cited ones were built by companies selling the private version of the same thing.

Start with the capability picture, because it is unusually clear-cut.

Enormous headroom in taste, none in compositionality

Three numbers, one conclusion

Visual Aesthetic Benchmark: the strongest frontier system identifies both the best and worst image correctly in 26.5% of tasks; human experts reach 68.9% (arXiv 2605.12684).

Design Crit: off-the-shelf VLM judges agree with designers about 54% of the time against a 50% chance floor; the human ceiling is 74.1%; a small trained head over a frozen encoder reaches 61.1% (contralabs.com/research/design-crit).

GenEval: the standard compositional text-to-image benchmark drifted to 17.7% absolute error against human judgement for current models and had to be replaced (arXiv 2512.16853).

Read those together. On the objective, checkable properties — is there one red cube and two blue spheres — the benchmarks have been solved hard enough that they now mismeasure the models they were built for. On aesthetic discrimination, the best system in the world is closer to chance than to a professional.

A 42-point gap between the frontier model and the human expert is not a research curiosity. It is the budget. Everything on the buyer side of this dossier — The generative-media labs, The creative tool layer — flows from the fact that the capability gap in taste is enormous, measured by two independent groups with different task formats, and shows no sign of closing on its own.

Saturation in this field is not one curve. Three distinct things are happening:

  1. Compositional and objective benchmarks are saturating and being replaced. GenEval → GenEval 2 is the clean case, with a quantified drift figure. VBench → VBench-2.0 is the same move in video.
  2. Aesthetic benchmarks are nowhere near saturation. 26.5% against 68.9%; 54% against a 50% floor. There is a headroom abundance.
  3. Arena-style leaderboards do not saturate; they drift. They are relative, so there is always a winner. That makes them durable marketing and weak products (The arena layer).

The design and UI benchmarks

BenchmarkBuilt byMethodScaleOpen?
UI-Bench (uibench.ai)Sam Jung, Agustin Garcinuno, Spencer Mateega — AfterQueryBlinded expert pairwise, TrueSkill with confidence intervals30 prompts, 300 sites, 194 expert judges, 4,047 pairwise matchesPrompts, framework and leaderboard open (arXiv 2508.20410)
UI Bench (ui-bench.dev)@vkiooz, independentFully automated — Cheerio static analysis 50%, axe-core 20%, Lighthouse 30%Landing pages, dashboards, mobileOpen (ui-bench.dev)
Design Arena"Intelligence"Crowdsourced pairwise across web/game/mobile, images, logos, slides, SVG, video, audio; captures agent traces, tool calls, re-promptsNot publishedLeaderboard public, data not (designarena.ai)
WebDev ArenaLMArena — now folded into Text Arena (Coding)Bradley-Terry over user votes, standardised React/TS/Tailwind promptNot publishedLeaderboard public (Epoch AI)
Human Creativity BenchmarkContra + MITPairwise + scalar + qualitative, phase-stratified28 evaluators / 13 countries, 5,940 pairwise, 5,940 scalar, 3,675 qualitativeCC BY 4.0 (arXiv 2606.30561)
Visual Aesthetic BenchmarkBakeLab, 17 authorsDirect ranking (best/worst) over matched subject matter400 tasks, 1,195 images, 10 expert judges per taskCC BY-NC-ND 4.0 (arXiv 2605.12684)
Design CritContra × Lica WorldCriteria-resolved preference; trained pairwise-difference head10 designers, 9 criteria, 1,600 ratings/criterionWrite-up public, data not (contralabs.com)

Two things to take from the table beyond the numbers. UI-Bench is the most rigorous public human-judged design benchmark in existence and it uses TrueSkill rather than Bradley-Terry specifically to get calibrated confidence intervals rather than point estimates — a methodological choice worth copying. And the two most useful design benchmarks disclose radically different things about their panels: VAB gives 10 expert judges per task and a controlled study design; UI-Bench names 194 experts with a role breakdown but states only that they were "hand selected and invited," thanking them for "generously contributing their time and judgment," which reads as unpaid (arXiv 2508.20410).

The image and video benchmarks

  • GenEval — compositional text-to-image, object-detection-based. Explicitly drifted. GenEval 2 documents benchmark drift of up to 17.7% absolute error against human judgement and introduces Soft-TIFA, decomposing into visual primitives (arXiv 2512.16853).
  • T2I-CompBench / ++ — compositional attributes, relationships, complex compositions (NeurIPS 2023 D&B). Academic, open, largely saturated on the simpler splits [WEAK].
  • GenAI-Bench — compositional text-to-visual generation with VQAScore-family automatic metrics; used as a transfer target by HPSv3++, which reports 76.3% (arXiv 2606.14657).
  • HEIM — Stanford CRFM's Holistic Evaluation of Text-to-Image Models; 12 aspects including aesthetics, originality, toxicity and fairness. Academic, open, and the intellectual ancestor of every multi-axis rubric in this market.
  • VBench / VBench-2.0 — Vchitect (arXiv 2503.21755). VBench-2.0 moves explicitly from "superficial faithfulness" to "intrinsic faithfulness": Human Fidelity, Controllability, Creativity, Physics, Commonsense. Extensive human annotation used to validate the automatic metrics — with no annotator count given.
  • EvalCrafter — text-to-video across visual quality, content quality, motion quality and text-video alignment, aggregated and calibrated to human opinion.
  • MJ-Bench — evaluates the judges, not the generators: alignment, safety, image quality, bias. Its key finding is directly monetisable for a critique-first vendor: natural-language Likert feedback from VLM judges is more accurate and stable than numerical scoring (arXiv 2407.04842). The written rationale beats the number.
"TASTE" as a named benchmark could not be resolved

Searches for a canonical benchmark under the name TASTE return the Visual Aesthetic Benchmark and unrelated work. The Contra × Lica dataset is referred to as TASTE in its own paper (arXiv 2605.20731) and as Design Crit on Contra's site, which may be the whole explanation — or the name may attach to something private or very new. [UNVERIFIED] Do not cite "the TASTE benchmark" as though it were an established public instrument.

Which ones are vendor-captured

This is the finding that should change a positioning plan.

UI-Bench was built by AfterQuery — "a research lab curating high-quality, human-generated data for AI foundational models" (arXiv 2508.20410 author affiliations; StartupHub). That is the same play as Contra Labs' Human Creativity Benchmark: an expert-data company publishes an open, rigorous, human-judged benchmark in its own vertical, uses it to establish methodological authority, and sells the private version.

AfterQuery ran that play in the design vertical roughly a year before Contra Labs launched. Any positioning that treats "publish a design benchmark to establish credibility" as novel is a year late in this specific niche — see AfterQuery and UI-Bench and Publishing the benchmark.

Sorting the landscape by who owns the scorer:

BenchmarkAuthor's commercial positionRead it as
UI-BenchAfterQuery — expert-data vendorVendor-published, methodologically strong
HCB / Design CritContra Labs — direct competitorVendor-published, and the source of the [[phase-decomposition|phase finding]]
Design ArenaIntelligence — sells the eval dataVendor-owned, data not released
WebDev ArenaLMArena — sells enterprise accessVendor-owned, data retained
VABBakeLab, academic, 17 authorsIndependent
GenEval / GenEval 2, T2I-CompBench, HEIM, VBench, EvalCrafter, MJ-BenchAcademic groupsIndependent

I note without asserting causation that the top-ranked tool on UI-Bench, Orchids, holds the largest margin over the field. There is no evidence of any relationship between AfterQuery and Orchids and none is alleged. But the general point holds for the whole category: when the benchmark author is a commercial party in the same market, the benchmark is not neutral infrastructure, and buyers increasingly know it.

The counter-move is governance, and it is cheap: an independent methodological advisory board, published conflict-of-interest disclosures, and pre-registered evaluation protocols. No US vendor in this market currently offers any of the three, and it is exactly what a European institutional buyer responds to (The one right a US competitor cannot hold).

What this means for a new entrant

Do not publish a benchmark to establish credibility. Publish a benchmark to establish a method. The credibility play is saturated; the methodological ground is not. The two open positions are the reliability-per-axis-and-phase work in Phase decomposition and the panel-sizing question in The oracle problem — neither of which any competitor has published, both of which are cheap, and both of which produce a number a buyer can use rather than a leaderboard they will forget.

And remember what a published benchmark costs you. Every published benchmark is a spent asset — the moment a de-novo brief set is public, it can leak into pre-training and its value as an uncontaminated test collapses. The resolution is the two-tier corpus in Eight units and one cost anchor: a public tranche sized for citation and inbound, and a permanently sealed tranche that is the actual product.

What the benchmark literature does not report

Annotator provenance is the least-documented dimension of the entire field. VBench-2.0 says "extensive human annotations" with no number. MJ-Bench discloses no annotator count. AgentNet discloses neither recruitment nor pay. UI-Bench names 194 experts and discloses no compensation.

That gap is closing for regulatory reasons rather than scientific ones — documented data provenance is becoming a compliance artefact under the AI Act (The one right a US competitor cannot hold) — which means a vendor that publishes panel composition, recruitment channel, pay and per-axis reliability converts a cost into a differentiator, and does so before it is compulsory.