Miju Labs

The security dossier

Telling a good designer from a confident one

There is a published pass mark: human experts hit 68.9% on the Visual Aesthetic Benchmark where the best frontier model manages 26.5%, and direct ranking yields substantially higher inter-annotator agreement than score-derived ranking. Collect rankings, screen at roughly 69% agreement with expert consensus, and treat peer nomination as a seeding mechanism rather than a scaling one.

high confidence8 minupdated 2026-08-30selection · work test · inter-rater agreement · nomination · portfolio · screening

Every supply business in this atlas eventually hits the same wall: the credential does not predict the work. In security you can set a machine-graded challenge. In medicine the state maintains the register. In creative work there is no register, the portfolio is a curated artefact, and confidence correlates with self-presentation rather than with judgement. So you need a work test — and unusually, someone has already published the pass mark.

The pass mark exists and it is 69%

The Visual Aesthetic Benchmark

400 tasks over 1,195 images across fine art, photography and illustration, with consensus labels from 10 independent expert judges per task. The strongest frontier system identifies both the best and the worst image correctly across three permutations in 26.5% of tasks. Human experts reach 68.9% (arXiv 2605.12684).

Two things fall out of that paper, and both are directly operational.

First, collect rankings, not scores. VAB's central methodological finding is that direct ranking by experts produced "substantially higher inter-annotator agreement on best- and worst-image labels" than rankings derived from individual scores (arXiv 2605.12684). Asking a designer to score an image 1–7 and then sorting the scores throws away agreement that asking them to rank the images directly retains. That is cheaper per unit of reliable signal, and it costs nothing to adopt.

Second, 68.9% is your screening threshold. It is the measured performance of the population you are claiming to sell. An applicant who cannot reach roughly 69% agreement with expert consensus on a held-out set is not an expert, whatever their portfolio says. That is a work test with a published, defensible, externally-sourced pass mark — the cleanest screening instrument available anywhere in creative supply.

Note what the same number does on the demand side: 68.9% is also a ceiling, and a mediocre one. Two experts agree about which of two images is better roughly two times in three. That is the whole subject of The oracle problem, and it means the screen has to be run against a consensus label from a panel, never against one senior person's opinion.

Make the work test the work

Contra's three-stage bar is portfolio review of a Contra profile, a video interview compressed for some roles to a 30-second intro, and a skills assessment that is "evaluating AI-generated brand assets and providing structured feedback" — advertised at roughly ten minutes of applicant time (per Contra Labs, sourced to help.contra.com and contralabs.com/jobs).

The design there is the point. The work test is the work. Every minute spent screening produces a labelled data point against a held-out set whose consensus you already know. That converts screening from a cost centre into partially-paid-for production, and it is the single most efficient piece of operational design in the competitive set. Copy it.

The corollary is that you must pay for the test. An unpaid assessment that produces usable data is the exact grievance the Sora signatories organised around — "we received access with the promise to be early testers… we are being lured into 'art washing'" — and it costs you the top of the funnel to save a few hundred dollars.

Order the stages by cost and by discriminating power:

StageWhat it screens forMarginal costDiscriminating power
Attitudinal pre-screenWhether they will do the work at all~$0High — the 81%/80% split is a sampling artefact you can sort on
Award / jury credential lookupPeer judgement of fitness to judge~$0 for jurorsHigh, but biased and small
Portfolio reviewCraft, not judgementHuman minutes, unbenchmarkedMedium
Ranked-choice work test vs. consensusJudgement, directlyPaid, and it yields dataHighest
Ongoing inter-rater agreementDrift, fatigue, collusionFree, from production dataHigh, and it never stops

The attitudinal question goes first, before any skills assessment, and it is allowed to disqualify. That is not squeamishness; it is arithmetic. Exeter measured 81% of designers saying AI dulls creativity while Contra's paid, self-selected panel answered 80% the other way (Dezeen; arXiv 2606.30561). Sorting on sentiment is free and it is the highest-leverage filter you have.

Peer nomination is a seeding mechanism, not a scaling one

Taste Labs' investor describes the mechanism explicitly: Taste "seed[s] the community with people vetted for strong aesthetic judgment, then allow[s] those tastemakers to nominate others," producing "a community where taste is socially validated — which is, if you think about it, how taste has always propagated," with the stated goal of "increasing the crowd's leverage without diluting its taste" (Amplify Partners).

It is a genuinely good idea. It solves the hardest problem in expert data — how do you know the expert is actually good — without a scaled ops team, and it produces a network with social cost to defection.

But the public front door contradicts the private mechanism. The Taste Makers portal is open application: "Apply to the opportunity that fits your craft," then "Take a short test to join the project. Pass, and you're placed on a project. If not, you will be welcomed into the community to be matched when the right project launches" (Taste Makers portal). Open application plus a work test, with nomination reserved for seeding and the top tier [WEAK — I am reconciling an investor's description with a product page, not reading an internal process].

Nomination gets you your first 200 people and your credibility. A work test with a published pass mark gets you your next 2,000. Treat the two as different instruments and cap nomination depth, because an uncapped nomination graph drifts toward a single school of taste — which is a defect in a dataset whose commercial value partly lies in mapping legitimate divergence. See Taste Labs, in full for what else Taste is building.

Awards as a prior, with the bias priced in

An ADC or D&AD jury seat is the strongest single credential available, because it is a peer judgement about someone's fitness to judge — which is literally the job. A Pencil or a TDC win is the second strongest. Roughly 800–1,000 jurors are named publicly every year with title and country, at zero cost to identify (the register argument).

Price the bias in, though. 80% of creative professionals entered no award in the past year, only 12% think awards help their careers, and 35% call them too expensive (Creative Boom); ADC freelancer entries run $75–$220 each (ADC). A screen built only on awards produces an agency-skewed, career-motivated panel and misses the independent with the most bench time. Use the award record as a prior that shifts the pass threshold, never as the gate.

Steal the Awwwards scoring machinery

Awwwards runs the most transparent scoring system in the creative industry and it is directly reusable. Every submitted site is scored by a minimum of 18 jury members; the system automatically discards the three scores furthest from the average; voting runs five days; only "Professional users" who have passed validation may vote in the community layer, explicitly to defeat sockpuppets. The rubric is fixed and weighted: Design 40%, Usability 30%, Creativity 20%, Content 10% (Awwwards).

Three things there are worth taking outright: an 18-rater panel per item, trimmed-mean outlier rejection rather than simple averaging, and a validated-professional gate on who may vote at all. The 40/30/20/10 weighting is a ready-made industry-recognised rubric, and it maps closely onto Contra's published five-axis rubric — visual quality and aesthetics, prompt adherence, originality, utility, motion realism (Contra Labs).

Agreement is an ongoing screen, not a one-time gate

The mistake to avoid is treating selection as an event. A rater who passed at 71% in month one can drift, fatigue, learn to game the consensus, or quietly start using a model to answer. The defence is that inter-rater agreement is a production statistic, computed continuously and reported per rater, per axis and per phase.

Contra reports Krippendorff's α and Friedman test p-values by domain and phase across 28 evaluators from 13 countries, producing 5,940 pairwise judgements, 5,940 scalar ratings and 3,675 qualitative rationales (arXiv 2606.30561). Krippendorff's α is the right statistic, and publishing it is the right move — it is what makes a taste dataset auditable to a buyer, and it is the artefact that maps onto AI Act documentation obligations (The one right a US competitor cannot hold).

Two design consequences. Agreement is axis-dependent — HCB finds it highest on prompt adherence and lowest on visual appeal — so a per-rater threshold has to be axis-relative or you will fire your best colourists for disagreeing about beauty. And agreement is phase-dependent, rising monotonically from ideation to refinement (Phase decomposition), so the same rater will look more reliable on late-stage work regardless of skill.

The stack that falls out, end to end: seed from juror rosters by name; propagate by capped nomination; open the funnel through Sidebar and ADPList (Where the good ones actually are); pre-screen on attitude; run a paid ranked-choice work test against expert consensus with a ~69% pass mark; then panel at ~18 raters where budget allows, with trimmed-mean rejection and α reported per axis. Getting the first cohort is the hard part; everything after it is arithmetic.

What nobody has published

There is no benchmark anywhere for the cost or reviewer-hours of professional design portfolio review at scale. Every operator has moved to machine-scored work tests with a human panel above them, but none has published what the human layer costs — so the trade-off between portfolio review and work test is asserted rather than measured.

Worse for planning: no study establishes, for professional design evaluation, the panel size needed to reach a stated reliability target per axis. Practice ranges from Contra's "3+ evaluators" to VAB's 10 to Awwwards' 18, with no published basis for choosing. That gap sets your entire cost structure and it is the subject of The oracle problem.