Miju Labs

The security dossier

Ninety days in taste

A sequence with a disproof attached to every step: seed from the published award juror rosters but never treat them as the funnel, ask the attitudinal question before the skills test, screen at roughly 69% agreement using rankings rather than scores, put the §40a royalty clause in the agreement before the first German contributor signs, split the verification and presentation ledgers in month one, and run the synthetic-versus-human ablation yourself rather than asking a buyer to run it.

medium confidence10 minupdated 2026-08-30execution · sequencing · recruiting · work test · ablation · metrics

Every step below is framed by what it is meant to disprove, because a ninety-day plan that cannot fail is a marketing document. The order matters more than the speed: three of these steps are cheap to do early and very expensive to retrofit.

Days 0–7: the two decisions that cannot be taken later

Incorporate so the EU entity is the maker. The party that takes the initiative and assumes the risk of investment is the one that holds the sui generis database right, and a cost-plus EU service subsidiary of a US parent may hold nothing at all. This is a one-hour conversation at incorporation and a restructuring afterwards.

Open three cost centres before you spend anything: verification, presentation, and obtaining. Adjudication, calibration, gold re-tests, rater exclusion and reliability computation go in the first. Schema, ontology, normalisation, indexing, API and documentation go in the second. Curation of award archives and published work goes in the third. Everything else is residual. BHB v William Hill excludes investment in creating data, and a litigator reconstructing the split from a general ledger three years later will lose.

What this disproves: nothing. It is the one step with no hypothesis attached, which is exactly why it gets deferred. Do it in week one because the cost of doing it in year two is the asset itself.

Days 0–14: the agreement, with the royalty clause already in it

Draft the The clause that expires your corpus in year ten before recruiting, not alongside it. The clause that has to be there from the first version is an ongoing royalty or corpus-participation component, however small — it takes the German grant outside §40a's lump-sum trigger, it is the cheapest defence to a §32a claim, and it satisfies the French L131-4 default. Re-papering several hundred contributors later means re-contacting people who have no reason to help you.

Also in version one: exclusivity on the artefact and never on the person; the byline; decline rights; and the employer-and-client warranty with an onboarding attestation behind it.

What this disproves: that European contributor law is a drafting problem you can solve after product-market fit. If counsel comes back saying the royalty component makes the unit economics unworkable, you have learned something real about the cost structure in week two rather than in diligence.

Days 7–21: seed from the juror rosters, by name

Roughly 800–1,000 named senior creatives sit on D&AD, ADC, TDC, Motion Awards and Awwwards juries every year, published with names and employers, refreshed annually, at zero cost to identify (the register argument). Write to two hundred of them individually, with a rate, a byline and a named person signing the email.

Do not treat this list as the funnel. It is the seed. It is too small, too agency-skewed and too expensive to reach at volume — 80% of creative professionals entered no award last year, and a panel built only on award records misses the independent with the most bench time. Use the jurors to establish that the panel is credible, then propagate by capped peer nomination and open the volume tier through paid channels.

What this disproves: that senior credentialed creatives will not answer. A reply rate is measurable in ten days and it is the single earliest read on whether the supply thesis is alive. If fewer than one in ten named jurors replies to a personally addressed offer with a real rate, the sentiment problem is worse than the evidence suggests and everything downstream needs re-pricing.

Days 14–30: ask the attitudinal question before the skills test

Exeter's DIGIT Lab found 81% of designers believe AI dulls creativity; Contra's paid, self-selected panel answered 80% the other way (Dezeen; arXiv 2606.30561). Those are accurate measurements of two different populations, and the difference is selection.

So the first screening question is attitudinal, it comes before any skills assessment, and it is allowed to disqualify. It costs nothing and it is the highest-leverage filter available. Running it after the skills test wastes the expensive step on people who were never going to do the work.

What this disproves: that sentiment is not sortable. If the attitudinal pre-screen passes almost everyone, it is badly written; if it passes almost nobody in a population that self-selected into your inbox, the addressable supply is much smaller than the headcount figures imply.

Days 21–40: a paid work test, in rankings, at 69%

Collect rankings, not scores. The Visual Aesthetic Benchmark's central methodological finding is that direct ranking by experts produced substantially higher inter-annotator agreement on best- and worst-image labels than rankings derived from individual scores (arXiv 2605.12684). It is cheaper per unit of reliable signal and costs nothing to adopt.

Screen at roughly 69% agreement with expert consensus on a held-out set. That is the measured human-expert figure on VAB, against 26.5% for the strongest frontier model — a published, externally sourced, defensible pass mark, and the cleanest screening instrument in creative supply (Telling a good designer from a confident one). Pay for the test. An unpaid assessment that produces usable data is precisely the grievance the Sora signatories organised around.

What this disproves: that a portfolio predicts judgement. The measurable output is the correlation between portfolio-review rank and work-test score across your first hundred applicants. If that correlation is strong, portfolio review is a cheap pre-filter. If it is near zero — which is what the credential literature would predict — you have just justified skipping the most labour-intensive step in the funnel, and that is worth more than the applicants.

Days 30–55: de-novo briefs only, and a first corpus with a reliability number on it

Write the brief bank yourself. Never use client work, in any product line, including trajectories. A recording of a real client project captures the brief, unreleased brand assets and possibly customer data, the contributor cannot license what they do not own, and the injured party is a client with no contract with you. De-novo briefs cost realism; pay the cost. They are also the only tranche that is provably uncontaminated by anyone's pre-training corpus, which is the property a buyer cannot verify any other way.

Then run a pilot corpus and report Krippendorff's α per axis and per phase, not as a headline but as a deliverable. Agreement is highest on prompt adherence and lowest on visual appeal, and it rises monotonically from ideation to refinement — so a single corpus-level reliability number is close to meaningless.

What this disproves: that de-novo briefs discriminate between models. A brief that every model handles identically is a wasted brief. The measurable output is the spread of model win rates per brief; briefs with no spread go in the bin and the ones that separate models become the sealed suite.

Days 40–70: the panel-size study, on data you already have

The single most valuable piece of first-party research available, and it is a re-analysis of the pilot rather than a new collection. No study establishes, for professional design evaluation, the panel size needed to reach a stated reliability target per axis. Practice ranges from Contra's "3+ evaluators" to VAB's 10 to Awwwards' 18, with no published basis for choosing between them (The oracle problem).

Subsample your own pilot panel at n = 3, 5, 7, 10, 15 and plot achieved reliability per axis. That curve sets your entire cost structure, it is publishable, no competitor has published it, and it costs a week of one person's time.

What this disproves: uniform panel depth. If the curve is flat above n = 5 on prompt adherence and still climbing at n = 15 on visual appeal — which is what the axis-agreement ordering predicts — then uniform panels overspend on the easy axes and underspend on the hard ones, and varying panel depth by axis becomes a defensible claim no competitor is making.

Days 55–85: run the ablation yourself

AgentTrek synthesises GUI trajectories from web tutorials at "a cost of just $0.55 per high-quality trajectory without human annotators" (arXiv 2412.09605). A narrated professional session costs $400–$1,500 all-in (The trajectory moat). That is roughly a thousandfold ratio, and the entire commercial case rests on the claim that the narrated thought and the preferred-execution path carry information the synthetic pipeline cannot recover.

Budget to prove it yourself. A buyer asked to run that ablation will not run it; they will decline the meeting. Train a small model with and without your trajectories, on a task where the difference should show, and publish the delta.

What this disproves: the moat. This is the step most likely to return a bad answer, which is exactly why it belongs inside ninety days rather than after a seed round has been raised on the assumption. If the delta is small, the trajectory product is not the wedge and the preference and critique units are — a survivable finding in month three and a fatal one in month eighteen.

Days 70–90: two buyer conversations, with the paperwork attached

Not a sales process — the buy side cannot be mapped by scraping and the routes in are relationships (Seventy people and one missing bridge). Two conversations, each with a written spec, a reliability certificate and the provenance package attached, are worth more than twenty introductions.

What this disproves: that documentation is worth paying for. If two buyers in a row treat the consent ledger and the Article 53 template language as table stakes rather than as value, the compliance premium is not a premium and the pricing model needs rebuilding.

The numbers that mean continue or stop

DayMeasureContinue ifStop or rethink if
30Reply rate from 200 named jurors≥10% reply, ≥20 in conversation<5%, or replies are hostile rather than uninterested
30Attitudinal pre-screen pass rate40–70% of a self-selected inbound poolAbove 90% (question is not discriminating) or below 20% (supply is smaller than modelled)
60Work-test pass rate at ~69%25–50% of attitudinally screened applicantsUnder 10% — the screen is unpassable, or the consensus set is wrong
60Krippendorff's α on the pilot, per axisα materially above chance on at least the craft axesNear-zero α on every axis — there is no convergent core to sell
60Verification + presentation share of spend, in their own ledgersA defensible fraction, documented and datedEverything still in one bucket
90Ablation delta, human versus synthetic trajectoriesA measurable, reportable improvementNo delta — pivot the product line, not the company
90Buyer conversations with a written spec returned2, with a named technical contactZero, after warm routes were used
The one that decides it

The 90-day ablation. Everything else on this page is recoverable. Recruiting can be re-run, the work test can be re-tuned, the ledger can be argued about. But if a model trained with narrated professional trajectories is indistinguishable from one trained on $0.55 synthetic ones, then the most defensible unit in the taxonomy is not defensible, and the business is a preference-data business competing on price with the arena layer. Find that out in month three, on your own money, in your own lab.

What is deliberately absent from this plan: a public benchmark launch. Publishing one is the obvious move and it is a year late in this specific niche — AfterQuery ran the play in design with UI-Bench roughly a year before Contra launched the Human Creativity Benchmark. The publishable artefact that is not already taken is the panel-size curve from days 40–70, and it is better marketing precisely because it is a method rather than a leaderboard. The full list of what remains unknown, ranked by how much it moves the decision, is in What the taste dossier could not establish.