Miju Labs

The security dossier

What this plan could not establish

Ranked by how much the answer moves the build. The top item is that the technical stack was not researched at all — environment interfaces, capture APIs, how FigmaTrace was actually made, and what an aesthetic reward model costs to train — which means no environment item can be costed in front of a buyer yet. Then four smaller questions, each with the specific action that settles it.

high confidence6 minupdated 2026-08-30gaps · open questions · build stack · diligence · environments · model cards

Ranked by how much the answer changes what gets built. Each carries the action that would settle it, not a plan to think about it further.

1. The technical build stack was not researched

The largest hole in this section

The evidence base behind these seven pages is a publishing and artefact audit — APIs, citation graphs, licences, download counts. None of the engineering questions were researched at all. Specifically unexamined:

  • Prime Intellect's Environments Hub and the verifiers interface — what shape a verifier has to implement, and whether a design environment fits it.
  • harbor and Inspect conventions — AfterQuery's harbor is confirmed to be the official harness for Terminal-Bench 2.0 and confirmed to contain no creative or design environments; what it would take to add one was not examined.
  • Figma Plugin API and Photoshop UXP capture limits — what can actually be recorded from inside each tool, at what fidelity, and what has to be captured from the screen instead.
  • How FigmaTrace was actually constructed — the paper describes a phase-based video-to-trajectory conversion over 200+ hours of expert video; the pipeline, the annotation cost and the failure rate were not established.
  • What an aesthetic reward model costs to train — compute, engineering time and data volume for a checkpoint of the taste-scorer class.
  • Narration-alignment cost per captured hour — aligning spoken narration to steps is the single largest post-processing line in a narrated-trajectory pipeline and no figure for it exists here or anywhere public.

Why it ranks first. Two items on What nobody has built — the RL environment and deeper trajectory capture — cannot be costed without it, and one of them is the largest prize on that list. Putting a number on either in front of a buyer before this is researched is guessing in public.

The action. A one-week technical scoping pass, run by an engineer rather than an analyst: read the verifiers and harbor interfaces and write a throwaway design environment against each; build a Figma plugin and a Photoshop UXP script that log one real session end to end; read the FigmaTrace paper's method section against its released repo; and price a reward-model training run on the TASTE data, which is MIT-licensed and already public. Everything needed is downloadable today.

2. Whether Design Arena has released preference data privately

What is unknown. designarena exists as a paid Hugging Face "team" org with three members, seven followers and a tagline naming RL, alignment and human preference — and zero datasets, zero models, zero Spaces, zero papers. Somebody is paying for a distribution surface that has never distributed anything.

Two readings fit equally well. A plan provisioned and not yet executed — the org was set up in anticipation of a release that has not happened. Or private distribution, where the org exists to host gated repos for named buyers and nothing is public by design. [WEAK] either way; the record does not separate them.

Why it matters. If Design Arena is already selling crowd preference data to labs, then the crowd-versus-panel argument at the centre of The taste read is being decided commercially while everyone argues it methodologically, and the head-to-head study at What nobody has built becomes more urgent rather than less.

The action. Ask them. A direct question to a company that has given two Hacker News launch posts and taken a $7.9M seed from Index is cheap, and a refusal is itself informative. Failing that, check whether any gated repo under the org resolves to a 401 rather than a 404 — a gated repo and a non-existent one are distinguishable.

3. Whether DesignPref's 12k comparisons were ever released

What is unknown. TASTE cites DesignPref at Krippendorff's α = 0.25 on a single overall label, from 20 designers and 12,000 comparisons — the only directly comparable professional-designer preference corpus in the literature, and the benchmark against which TASTE's own α ≈ 0.19 should be read. TASTE states the data had not been released at its submission time. Current status was not verified.

Why it matters. If DesignPref is public, it is a free second dataset for the panel-size sweep at Rebuild the Human Creativity Benchmark — a re-analysis at zero collection cost on 12,000 comparisons from 20 designers, which is a deeper panel than anything else available. If it is not, that changes nothing about the plan but removes a shortcut.

The action. One search of the authors' pages and one email. Yi-Hao Peng is named in the acknowledgements of Taste Labs' Requests for Research and DesignPref is on its reading list, so the route is short.

4. What Contra paid its 28–31 HCB evaluators

What is unknown. The HCB paper discloses only the $350 formative-survey fee paid to each of 50 respondents in November 2025, and never states what the benchmark evaluators were paid. By contrast the TASTE paper discloses precisely — "a flat project fee that averaged approximately $90 per hour", 13 and 16 hours per cohort. The flagship paper omits the number its collaborator published.

Why it matters. Less than it looks for the build, more than it looks for the pitch. The $90/hr anchor already exists and is used throughout Eight units and one cost anchor and the costing at Rebuild the Human Creativity Benchmark; a second data point would tighten the band but not move it. What the omission does is create the market's most answerable disclosure gap, at a company whose recruiting proposition is built on paying creatives well. A competitor that publishes panel pay as a line item in every release converts an omission into a differentiator — and nobody in this market publishes it.

The action. No research settles it; only a disclosure would. Treat it as a question to ask if anyone from Contra is in the room, and as a standard to adopt regardless.

5. Whether any of these artefacts has ever appeared in a lab's model card

What is unknown, and what the record does show. No model card, system card or lab publication found in this pull references Contra Labs, Lica, purvanshi/TASTE or AfterQuery/ui-bench. Zero models on Hugging Face declare training or fine-tuning on any of them; FinanceQA has exactly one declared dependent. That is negative search evidence and is [WEAK] as proof of absence, but it is consistent across the citation graph, the HF model index and the press record.

Why it matters most of all. It is the only measure of whether any of this ever reached a training run. Citations measure whether researchers read you. Downloads measure whether anyone clicked. A model card mention measures whether a lab depended on you, and on the current record nothing in this market has ever cleared that bar. If that stays true after a proper search, it is the strongest single argument that the artefact is customer acquisition and the panel is the business — which is the position at What to build first and The oracle problem.

The action. A systematic sweep of published model and system cards from the major labs for named external data suppliers of any kind, not just creative ones. If external suppliers are named routinely and creative ones never are, that is a finding about this market. If suppliers are never named at all, it is a finding about disclosure norms, and the AI Act Article 53 template is about to change it (Sell the paperwork with the data).

Also open, and smaller

The 48% discrepancy. HCB's paper describes 15,555 judgments; the public dataset holds 7,537, with an evaluator count that differs by three in the wrong direction. Neither document explains it. There may be an entirely ordinary reason. Nobody has asked.

Any price, anywhere. Not one organisation in the record publishes a rate for a preference dataset, a trajectory hour or an evaluation engagement — which means the first credible published rate card sets the market (Eight units and one cost anchor).

Whether a publication has ever produced a customer conversation, for anyone. No company in this set publishes a case study tying a paper to a deal. Sacra's characterisation of AfterQuery's benchmarks as driving "sales enablement" is an analyst's read, not a disclosure. [UNVERIFIED]

The wider set of unresolved questions for the whole lane is at What the taste dossier could not establish; the sequence that would close several of them in ninety days is at Ninety days in taste.