Miju Labs

The security dossier

Eight units and one cost anchor

The competitor published its own cost base: $90/hr, 10 designers, 14,400 comparisons, about $13,000 — roughly $0.91 per pairwise judgement. Extrapolated, a frontier-grade 1.17M-comparison corpus costs about $1M in direct labour, which is a build-versus-buy decision a lab makes in one meeting. Everything defensible in the taxonomy sits above that line.

medium confidence10 minupdated 2026-08-30artefacts · pricing · unit economics · preference data · rubrics · eval suites

Start with the number the competitor published about itself, because it governs the whole page.

The cost anchor

From the Contra × Lica TASTE paper: 10 professional graphic designers in two disjoint cohorts of five, 4–10 years' experience, screened by portfolio review. 80 prompts × 5 designers × 4 generators = 1,600 ratings per sub-dimension, totalling 14,400 pairwise comparisons across nine criteria. "Designers were paid a flat project fee that averaged approximately $90 per hour." One cohort worked ~13 hours, the other ~16 (arXiv 2605.20731).

10 × ~14.5 hours × $90 ≈ $13,050 for 14,400 comparisons ≈ $0.91 per pairwise comparison.

Now extrapolate against the field, because that is what a buyer will do the moment you quote them.

Target corpusComparisonsDirect labour at $0.91
Contra's own HCB5,940~$5,400
Contra's TASTE / Design Crit14,400~$13,000
ImageReward137,000~$125,000
HPD v2798,090~$726,000
MPS / MHP918,315~$836,000
HPDv31,170,000~$1.06M
Arena's text-to-image board6,047,075~$5.5M — Arena paid roughly $0

A frontier-competitive professional aesthetic preference corpus costs about $1M in direct annotator labour. That is a fundable line item, not a moat. Any lab with $1M and a programme manager can build it, and several already have. It also means Contra's entire published corpus cost under $20,000 in designer time — their defensibility is not the data.

Three consequences run through everything below. Price comparison volume as a consumable, not as IP. Assume the buyer models the build cost before the first call. And put the margin story on the high-touch, low-volume, high-ASP units, because on volume the arena layer beats you by five orders of magnitude on cost.

The unit table

UnitWhat it isDirect costWho buysDefensibility
Pairwise comparison(prompt, A, B, winner, rater_id, axis)$1.50–$6 per judgement [WEAK]RL/post-training, eval teamsLowest — public format, public schema
Rubric-scored critiquePer-axis scalar + 100–400 words of reasoning$8–$25 each [WEAK]SFT / corrective-data, reward-model teamsMedium — the rubric is the asset
Multi-axis quality scoring5–9 scalar axes per output$3–$10 per output [WEAK]Reward-model teams, product QAMedium-low alone; high as a longitudinal panel
Phase-decomposed evaluationThe above, stratified by workflow stage+20–40% over flat scoringModel-improvement teams (diagnostic buy)High — it is a method, and methods travel
Process trajectoryScreen recording → state/action pairs + narrated thought$400–$1,500 each [WEAK]Computer-use / agent teamsHighest — cannot be scraped
De-novo brief + reference setOriginal prompt corpus, never web-published$150–$600 per brief [WEAK]Eval teams, benchmark buyersHigh while uncontaminated
Held-out private eval suiteDe-novo briefs + sealed judgements + hosted scoringSuite build $50k–$250k [WEAK]Frontier labs, regulated deployersHighest commercially — a subscription, not a file
Trained aesthetic reward modelPreference head over a frozen VLM encoderMarginal, given the corpusApp-layer companies, internal QAMedium — weights leak information, not the corpus

The cost column is the weakest thing on this page and it deserves the flag. No taste-data vendor publishes per-unit prices. These are back-derived from three disclosed anchors: Contra's "up to $100/hr" with briefs "typically 5 to 10 hours" (contralabs.com/jobs), the HCB study's ~15,000 professional judgements across 28 evaluators and 80 sessions (arXiv 2606.30561), and the $90/hr TASTE figure above. Treat them as planning bands, not quotes.

Pairwise comparisons — a consumable, and price it that way

The minimum viable unit: a prompt, two outputs, a winner, a rater identity and — in any competent implementation — an axis label, because "which is better" without an axis is the largest single source of unusable data in this market. Aggregation is Bradley-Terry (LMArena's method, Epoch AI) or TrueSkill, which UI-Bench chose because it yields calibrated confidence intervals rather than point estimates (arXiv 2508.20410).

Three structural reasons it is the least defensible unit. The format is fully public and the tooling is open source. The supply is substitutable, and a buyer cannot verify quality from the delivered file alone. And preference pairs have a short half-life — they are indexed to a model generation. HPSv3++ needed 212K new pairs over 185K prompts and 1.1M+ images specifically because reward models must be re-conditioned on current-generation outputs (arXiv 2606.14657).

Short half-life is not a criticism. Consumables produce recurring revenue. Just do not build the moat story on them.

Rubric-scored critiques — the rubric outlives the critique

Per-axis scalar plus free text. In HCB, 3,675 qualitative responses accompanied 5,940 scalar ratings (arXiv 2606.30561).

A scalar tells a model that it lost. A rationale tells it why, in the vocabulary of the discipline — the baseline grid breaks at the third module; the optical sizing is wrong at display weights; the hierarchy inverts because the eyebrow outweighs the headline. That vocabulary is exactly what a generalist crowd cannot produce and exactly what a frontier model's design critique currently lacks. There is external support for preferring text over numbers: MJ-Bench, which evaluates the judges rather than the generators, finds that natural-language Likert feedback from VLM judges is more accurate and stable than numerical scoring (arXiv 2407.04842). The written rationale outperforms the number.

But the defensible asset is the rubric, not the critiques written against it. A nine-axis rubric with anchored scale points and worked exemplars takes months and dozens of practitioner-hours to stabilise; the critiques are its renewable output. It is also the one piece of IP you author yourself rather than licensing from a contributor, which makes it the cleanest thing in the corpus to own (Copyright is the weakest thing you own). Version it, treat the version history as a trade secret, and licence it separately.

Multi-axis scoring — vary the panel depth by axis

Composition, typography, hierarchy, colour, craft, plus prompt adherence and utility. Contra's public rubric is five-axis with 3+ evaluators per output (contralabs.com/human-creativity-benchmark); its internal Design Crit work uses nine (contralabs.com/research/design-crit).

The empirical fact that should shape axis design: agreement is not uniform across axes, and the ordering is stable. HCB reports agreement highest on prompt adherence and lowest on visual appeal (arXiv 2606.30561) — the same ordering the academic aesthetics literature finds between domains, where Vessel et al. (2018, Cognition 179) found high shared taste for faces and landscapes but strong individual variation for architecture and artworks (Max Planck Neuroscience).

Different axes therefore need different numbers of raters to reach the same reliability. A uniform panel overspends on prompt adherence and underspends on visual appeal. A vendor that publishes per-axis reliability and varies panel depth accordingly is doing something no competitor advertises — and it is a claim you can make in a sales conversation without hedging.

Phase-decomposed evaluation

Covered properly in Phase decomposition. For the taxonomy: it costs 20–40% more than flat scoring in design and analysis overhead, it sells as a diagnostic rather than a dataset, and the durable part is not the headline inversion between models but the monotone rise in inter-rater agreement across ideation → mockup → refinement (Kendall's W 0.345 → 0.436 → 0.549 on ad images). That is a pricing schedule, not a finding about any particular model.

Process trajectories

The most defensible unit in the taxonomy and the subject of The trajectory moat. Screen-recorded professional sessions with narrated intent, structured into state-action pairs with execution_paths and preferred_execution. Planning band $400–$1,500 per usable trajectory against a synthetic alternative at $0.55 (arXiv 2412.09605) — a thousandfold ratio that you have to justify with an ablation you run yourself.

De-novo briefs and reference sets

A brief written for the purpose, never web-published, with a reference set of exemplar solutions. Cost is dominated by senior time: writing a brief that is realistic, unambiguous and discriminative between models is an art-direction task, not a writing task.

Two properties make it strategically important out of proportion to its revenue. It is the only unit that is provably uncontaminated — a buyer training on web-scale data has no way to know whether a public benchmark leaked into pre-training, and increasingly assumes it did. And it is the substrate for the held-out suite below.

Its defensibility decays the instant it is published, which is the tension inside the benchmark-as-marketing strategy that both Contra and AfterQuery run. The resolution is a two-tier corpus: a public tranche sized to generate citation and inbound, and a permanently sealed tranche that is never released and is the actual product.

Held-out private evaluation suites — the best business here

De-novo briefs, sealed expert reference judgements, a hosted scoring service, sold as a subscription with periodic refresh.

Why the sealed suite is the highest-margin unit

It converts a data sale — one-time, margin-compressing, resellable by the buyer — into a service relationship that recurs, cannot be resold, and whose switching cost grows with the length of the buyer's own score history. It is also the only unit where the absence of an oracle is an advantage: because ground truth is your panel's adjudicated judgement and your panel is held privately, the buyer cannot replicate your scores without you. In a domain with a real oracle, a held-out suite is just a test set the buyer could have built themselves (The oracle problem).

The security-evaluation market has already made this transition. The creative market has not, which is the opening — and it is the structural argument in What to build first.

Trained aesthetic reward models — occupy the axes, do not out-scale

Design Crit trains "a small pairwise-difference head on top of a frozen vision-language encoder, with no fine-tuning of the backbone." Off-the-shelf VLM judges reach ~54% agreement with designers against a 50% chance floor; the human ceiling is 74.1%; the trained head reaches 61.1% — roughly 45% of the way from the VLM to the human ceiling (contralabs.com/research/design-crit).

The general-purpose comparables are three orders of magnitude larger. HPSv3/HPSv3++ train on 212K preference pairs over 185K prompts and 1.1M+ images, reporting 79.1% aesthetic and 88.1% text-following accuracy on their own held-out data (arXiv 2606.14657).

The specialist play is not to out-scale the crowd datasets. It is to occupy the axes the crowd datasets collapse. HPSv3 has two dimensions. Design Crit has nine. That is the entire argument, and it is why the $1M extrapolation at the top of this page is a warning rather than a target: matching HPDv3 on volume buys you a corpus a lab could have commissioned itself, at a price it will notice.

What is not knowable about pricing

No vendor in this market publishes a per-unit price — not Contra, not Taste Labs, not the generalists. The $0.91 anchor rests on one disclosed figure, in one paper, for one study. Every reward-model paper in the academic line — ImageReward, HPD v2, HPSv3, MPS, VisionReward, VideoScore2 — discloses annotator counts and hides annotation budgets. AgentNet discloses neither recruitment nor pay (HF); UI-Bench names 194 experts and states only that they were "hand selected and invited," thanking them for "generously contributing their time and judgment," which reads as unpaid (arXiv 2508.20410).

Annotator provenance is the least-documented dimension of this entire field, at precisely the moment documented provenance becomes a compliance artefact (The one right a US competitor cannot hold). Publishing panel composition, recruitment channel, pay and per-axis reliability converts a cost into a differentiator — because nobody else does it.