Miju Labs

The security dossier

What nobody has built

Eight absences, ranked, each searched specifically and each with what it would take and what it would earn. Zero design reward-model weights exist from any of the six organisations examined; no design benchmark anywhere is held out; nobody has run designer-panel against crowd-vote on the same items though both companies whose business models depend on the answer have the data; and the trajectory position is now contested.

medium confidence12 minupdated 2026-08-30gaps · reward models · rl environments · held-out sets · provenance · trajectories

Each absence below was searched specifically against the Hugging Face API, arXiv, Semantic Scholar and the six organisations' own repositories. Two of the ones this list was expected to contain fell over on checking and are recorded as closed rather than quietly dropped. Ranked by what building it would earn.

1. The panel-size study

What is missing. How many designers per item an aesthetic ranking needs before it stabilises, and how that number varies by criterion. Practice runs from Contra's "3+ professional creative evaluators" through VAB's 10 expert judges per task to Awwwards' minimum of 18 jurors — a six-fold spread with no published derivation anywhere. Google's Forest vs Tree work says 1, 3 or 5 raters per item is "often insufficient" and that practitioners "often need more than 10", and it was not run on aesthetic judgment. TASTE built exactly the right machinery — Kendall's τ, pmax, Condorcet cycle rate against exact nulls — and ran it at a fixed R = 5.

This is now the subject of Rebuild the Human Creativity Benchmark, which has the design, the costing and the risks. It sits at number one here and is not repeated: read that page instead. The only thing to add is why it ranks above the seven that follow — it is the only item on this list that is a re-analysis rather than a new collection, which means it is the cheapest and fastest, and it sets the cost structure of everything else on the list.

2. A held-out creative eval set with a submission API

What is missing. Every design benchmark in this market is fully public. HCB's prompts, outputs and judgments are all in the CC-BY-4.0 repo. UI-Bench's prompts are on Hugging Face. GDB is a 17 GB public download. Design Arena's prompts are user-generated and never published at all. Nobody operates a held-out creative test set with a submission API — the standard structure in every mature benchmark category, from SWE-bench to Terminal-Bench to ARC-AGI.

What it takes. De-novo briefs written for the purpose and never web-published, at roughly $150–600 each of senior authoring time; sealed panel judgments against them; a submission endpoint and a scoring service; and the operational discipline never to leak the set, including to a customer who asks nicely. The technical component is a small web service. The hard part is organisational: the set is worth nothing the day it is published, and every incentive in a young company points at publishing it.

What it earns. Two things nothing else on this list earns. It is the only structure that keeps measuring something in a year, because a public benchmark is contaminated the moment a model trains on the internet that contains it. And it converts a data sale into a subscription whose switching cost grows with the length of the buyer's own score history — the highest-margin unit in the taxonomy and the shape argued at What to build first. Publish the dev split for citation; seal the mirror.

3. A longitudinal instrument

What is missing. Every creative benchmark in the record is a one-shot snapshot against a fixed model slate. Nobody re-runs the same instrument against successive model releases and publishes the trend line. Contra comes closest by sheer volume — 39 studies in 140 days — and cannot stack any of it, because each study uses a different panel, different prompts and a different rubric. Thirty-nine incompatible snapshots are not a series.

What it takes. Freezing four things and versioning them: the prompt set, the rubric, the panel roster and the aggregation method. Then re-running on a published cadence — quarterly, or on each frontier release — and reporting the same statistics each time with the panel's own drift measured alongside. The cost is not the re-run; it is the discipline of not improving the instrument between runs, and the retainer that keeps the panel stable enough to make consecutive scores comparable.

What it earns. A trend line is the only artefact in this market that gets more valuable with age, and it is the natural form of a subscription. It also produces the one thing a lab genuinely cannot generate internally on demand: a historical series measured by the same panel under the same protocol, against which its own model's progress can be read.

4. Open aesthetic reward-model weights

What is missing. Zero design-domain reward-model checkpoints are public from any of the six organisations examined. Lica's taste-scorer is the closest thing in existence: MIT-licensed inference code on a frozen Qwen3-VL-Embedding-2B backbone, reaching 0.611 pairwise accuracy against a 0.741 single-rater ceiling, with the analysis code and the data both published — and no trained checkpoint under purvanshi, lica-world or contralabs. Off-the-shelf VLM judges in the same evaluation land in a 4.4-point band from 0.499 to 0.543, none reaching 0.55.

What it takes. Almost nothing, given a corpus. A preference head over a frozen encoder is days of work, and the training run is marginal against the cost of the judgments it learns from. What it takes is the decision to publish the weights, and the acceptance that weights leak information about the corpus in a way a paper does not.

What it earns. The difference between an artefact people cite and an artefact people build on. A checkpoint gets imported into someone's training loop, appears in their model card, and shows up in the HF model index as a declared dependency — which is precisely the downstream signal that is currently zero across this entire cohort. It is the cheapest item on this list relative to what it earns, and the fact that nobody has shipped one is the most surprising absence in the record.

5. Designer panel versus crowd vote, on the same items

What is missing. Nobody has published professional-panel agreement against crowd-vote agreement on the same items, scored the same way. Design Arena's entire business rests on the claim that unpaid anonymous votes are good enough. Contra's rests on the claim that they are not. Both have the data to settle it and neither has published it.

What it takes. One item set, two rater populations, one aggregation method, one paper. The professional side is a panel you are building anyway. The crowd side can be bought on a general crowd platform for a few thousand dollars. The analysis is the same reliability apparatus as item 1, run twice.

What it earns. This is the single most commercially loaded question in the market, and whoever answers it credibly is quoted by everyone arguing either side of it for the next two years. It is also genuinely two-sided: if professionals converge no better than a screened crowd on the axes a buyer cares about, the premium-panel pitch is weaker than this whole dossier assumes, and it is better to learn that from your own study than from a buyer's. Pick-a-Pic's finding — that real users beat "expert annotators" who were the authors' friends and colleagues — is the version of this question that currently stands in for an answer, and it is not one (The taste read).

6. An open creative or design RL environment

What is missing. Nothing exists, anywhere. AfterQuery's harbor is a general agent-evaluation and RL-environment framework and contains no creative or design environments. Design Arena's agent-runner is a model-agnostic harness, not an environment. Contra publishes trajectories with no environment to replay them in. Lica's lica-bench is static-metric evaluation. Taste Labs names the problem — "what it takes to build high-fidelity environments of the digital world for models" — and has shipped nothing. No one has published a resettable, rewarded, design-tool environment an agent can be trained in.

Be precise about what is actually verifiable, because this is where enthusiasm outruns the evidence. Design splits cleanly into two halves, and only one of them can carry a reward function:

Checkable, so it can be an environmentContested, so it stays an eval with a human panel
Rendered-output comparison against a targetBrand fit and identity judgement
DOM and CSS assertionsIllustration quality
Contrast ratios and accessibility rulesMotion feel and timing
Brand-token adherence — does it use the declared palette, type scale, spacingWhether the concept is any good
Constraint satisfaction of the Arena-T2I-Hard kind — ~30 decomposed yes/no constraints per promptOriginality

Taste Labs' own disclosure is the strongest available evidence for this split: "A verifiable layout target lets a construction loop converge, while a contested aesthetic target does not: in our patent render-loop, 0/24 converged." An environment built on the left column is buildable and useful. An environment that claims to reward the right column is a reward-hacking generator, which is problem 7 on their own list.

What it earns. The largest prize on this page and the least certain. RL environments are the artefact class labs are visibly buying right now, and there is no creative one. But the honest position is that this plan cannot yet cost it.

The build stack for this has not been researched

Everything needed to scope an environment or a capture pipeline is outside the evidence base of this dossier: Figma Plugin API limits, Photoshop UXP capabilities, how FigmaTrace was actually constructed from 200+ hours of video, the interface shape a verifier has to implement to plug into an existing harness, and what training an aesthetic reward model actually costs in compute and engineering time. None of it was researched. Do not put a number on this item in front of a buyer until it has been. It is the top entry in What this plan could not establish.

What is missing. No schema, no spec, no reference implementation from any of the six organisations. Taste Labs asks for "techniques for detecting, attributing, and auditing design outputs in open-ended generative settings" and has published none. C2PA exists at the file level for media provenance and addresses none of the design-process attribution problem. The nearest thing to a de facto format is Contra's own trajectory schema — execution_paths, preferred_execution, narration-derived thought — published only as a dataset column layout and never as a specification.

What it takes. A written spec, a JSON schema, a validator, and one reference dataset that conforms to it. Weeks, not quarters. The expensive part is the ledger discipline behind it: dated consent records, licence chains, per-record provenance, a jurisdiction field.

What it earns. Timing is the whole argument. AI Act Article 53 enforcement powers went live on 2 August 2026, with fines to 3% of global turnover or €15m, and compliance is visibly patchy — Anthropic, Mistral and xAI substituted narrative prose for the Commission's template. A commissioned, consented, contract-backed corpus is the easiest line a lab will ever write into that summary, and no lab has published data-supplier requirements for creative content, which means the first credible published standard sets them. The full argument is Sell the paperwork with the data, with the drafting at The clause that expires your corpus in year ten and the property-right angle at The one right a US competitor cannot hold.

8. Deeper trajectory capture — now a contested position

What changed. This was the flagship gap and it closed ten days before this plan was commissioned. FigmaTrace (Patronus AI, arXiv 2608.21460, 19–20 August 2026) is CC-BY-4.0: 3,469 design trajectories from 200+ hours of captured expert video across 126 open-ended long-horizon tasks, 22.3 GB, with an open SFT model and a claim of parity with frontier systems on four out-of-distribution agentic GUI environments. 1,784 downloads in fourteen days. Contra's five trajectory repos hold 1,769 steps.

So the differentiator can no longer be volume. Three things remain genuinely unoccupied, and they are the ones to build:

Narrated intent. FigmaTrace's method is video-to-trajectory conversion; Contra's thought field is transcribed from the editor's spoken narration during the session, explicitly distinguished from model-synthesised rationales. Narration at scale is the harder capture and the one nobody has done in volume.

Multi-path execution. execution_paths plus preferred_execution encodes that expertise is partly route choice — a professional knows three ways to do a thing and reaches for one without thinking. A trajectory recording only the resulting state throws that away. No other corpus records it.

De-novo briefs. Trajectories captured during real client work carry the client's brief, unreleased assets and possibly their customer data in frame, and the contributor cannot license what they do not own (The trajectory moat). De-novo briefs cost more, lose the "this is real work" line, and are the only version that survives diligence and is provably uncontaminated.

The value ratio still has to be defended: a synthetic GUI trajectory costs $0.55 and a human one $400–1,500. No public ablation exists showing what narrated human creative trajectories buy over synthetic ones, and running that ablation yourself is the highest-leverage piece of first-party research on the product side.

Two absences that closed, recorded honestly

An open design-trajectory corpus — closed by FigmaTrace, above. Any plan asserting this gap is working from a map that is two weeks stale.

A reliability and panel-size studypartially closed by Google's Forest vs Tree, which establishes the general result and the simulator but not the domain-specific curve. The creative-domain version, per criterion, remains open and is item 1.

What could not be established at all is at What this plan could not establish; how any of these actually travels once built is at How an artefact travels.