Miju Labs

The security dossier

What publishing actually bought them

Thirty-nine studies, two papers, eight datasets and no GitHub organisation bought Contra Labs 5,133 all-time downloads and zero citations. Patronus out-published them on their own moat in fourteen days; AfterQuery's 148-row FinanceQA has 17 citations and a repo. Volume bought nothing, and the pattern that did work is small, sharp, cited and shipped with code.

high confidence8 minupdated 2026-08-30publishing · citations · downloads · scoreboard · figmatrace · distribution

Start with the scoreboard, because it settles an argument that would otherwise run for a quarter.

Contra Labs has eight Hugging Face datasets totalling roughly 5,100 all-time downloads and 24 likes, thirty-nine dated studies published between 8 April and 26 August 2026, two arXiv papers with zero citations each — the Human Creativity Benchmark (2606.30561) and TASTE (2605.20731) — and no GitHub organisation at all: github.com/orgs/contralabs and github.com/orgs/contra-hq both return 404. Their Hugging Face org carries zero models, zero Spaces, zero HF Papers and five followers.

Their best-performing artefact is premiere-video-editing-trajectories: four trajectories, 234 steps, 182 MB, 1,415 all-time downloads and 12 likes. It is a genuinely well-built schema — screenshot, narrated thought, structured tool_call, execution_paths, preferred_execution — and it is four trajectories.

The numbers in circulation were row counts, not downloads

Earlier notes on this desk reported Contra's top dataset at ~8,000 downloads with the rest at 15–800, and a striking likes-over-downloads pattern. Those were the row counts from the dataset-viewer size API — HCB has 8,012 rows, descript 803, photoshop 294 — read into the wrong column. Verified against the HF datasets API with expand[]=downloadsAllTime, the org total is 5,133 all-time downloads and 24 likes. The "attention without adoption" story is not supported: ratios run 0.004–0.014 likes per download across the whole org. Contra has neither attention nor adoption on this channel.

The scoreboard

OrgDatasetsAll-time downloadsLikesPapersCitationsGitHub stars
Contra Labs85,1332420no org exists
Lica World (now Gamma)4 incl. purvanshi/TASTE10,71415410 (8 self)13
AfterQuery614,54531632~38
Design Arena00~146
Taste Labs00no org
LMArena27 (three orgs)500,000+4,980 on one Space15+very high [WEAK]~47,000

Contra is the only organisation in that table with a zero-citation research corpus. TASTE has been public since 20 May 2026 and HCB since 29 June. Cadence was not the problem: 39 studies over 140 days is one every 3.6 days, and it produced roughly 1.3 dataset downloads per study published.

The release that took the moat in fourteen days

On 19–20 August 2026, Patronus AI released FigmaTrace: 3,469 design trajectories drawn from 200+ hours of captured expert video across 126 open-ended long-horizon tasks, 22.3 GB, CC-BY-4.0, with a paper (arXiv 2608.21460) and an open SFT model (Qwen3.8-27B-Figmatrace-SFT). In its first fortnight it took 1,784 downloads — more than Contra's best repo has taken in seven weeks, and roughly 35% of Contra's entire eight-dataset all-time total.

The moat comparison, in the same units

Contra's five trajectory repos hold 1,769 trajectory steps in total. FigmaTrace holds 3,469 complete trajectories. Contra published the format first, published it well, and was out-published on volume by a company that was not previously in this market, in a single release, ten days before this plan was commissioned.

That does not make the trajectory asset worthless — narrated intent, multi-path execution and de-novo briefs are all still open, and FigmaTrace's method is video-to-trajectory conversion rather than live narrated capture. It does mean the "nobody has published an open design-trajectory corpus" position is gone, and any plan that rests on being first to volume there is a plan written against a stale map.

The counter-example: small, sharp, cited

AfterQuery ran the opposite strategy and it worked.

ArtefactSizeDownloads (all-time)CitationsRepo?
AfterQuery/FinanceQA148 rows6,31117 (1 influential)FinanceQA, 9★
AfterQuery/vader2,659 files, 174 vulnerabilities5,0408vader, 11★
AfterQuery/ui-bench30 briefs, 5 categories3952framework open
AfterQuery/App-Bench6 rows627appbench.ai-docs, 1★
AfterQuery/MCP-Universe2 rows549

FinanceQA is 148 rows and has twelve times Contra's best repo's downloads and seventeen citations. VADER has eight. UI-Bench — the closest direct comparable to HCB, and the better-built one, with 4,047 blinded expert pairwise matches and TrueSkill confidence intervals — has two, and shipped eleven months earlier (AfterQuery and UI-Bench).

The structural difference is not size and it is not marketing spend. Every cited AfterQuery artefact has a GitHub repository beside it. Every uncited Contra artefact does not. AfterQuery published a domain benchmark roughly every four months for eighteen months, each naming specific commercial tools and showing them failing, each with a live leaderboard on its own domain — uibench.ai, marketbench.ai, ide-bench.com, appbench.ai — and each runnable by someone who was not them. Over the same period the company went from YC W25 to $100M annualised and a $3.2B valuation (Sacra). Causation is not established and should not be claimed. The correlation between shipping code with the data and being cited is not subtle.

Lica: volume moves downloads and does not move citations

lica-world/GDB is 33,886 rows across 40 task configs at 17.0 GB — roughly a hundred times Contra's total published bytes — and it has 4,615 all-time downloads, three times Contra's best. Volume does buy downloads.

It does not buy standing. The LICA paper (2603.16098) has 5 citations, of which 4 are Lica's own papers; the single external citer is Does Synthetic Layered Design Data Benefit Layered Design Decomposition?. GDB (2604.04192) has 4 citations, 3 self. Say that plainly rather than quoting the headline: the Lica corpus has ten citations and eight of them are its own. The one peer-reviewed venue in the entire Contra/Lica corpus is a workshop — ICML 2026 Human-AI Co-Creativity — for the design-video metrics paper, which has zero citations.

There is a second lesson buried in Lica's card copy. lica-world/lica-dataset publicly contains 1,148 compositions, described as a curated sample of the 1,550,244 the paper claims. The "1.5M layered compositions" that anchors Lica's whole positioning is a paper claim; sub-0.1% of it is downloadable.

The scale reference, so nobody mistakes the ceiling

LMArena's lmarena-ai/leaderboard-dataset holds 2,277,369 rows across 21 configs and has 126,168 all-time downloads, refreshed continuously — it was last modified on the day the inventory was pulled. Its arena-leaderboard Space has 4,980 likes, two orders of magnitude more than any other artefact in this market. Its GitHub is about 47,000 stars, of which FastChat alone is 40,000.

That is the asset class Contra is implicitly comparing itself to, and it is not a data asset. It is nine years of accumulated open-source goodwill, and it is why LMArena can charge for evaluation and Contra cannot. Nothing in this plan should aim at it. It is here as a ruler.

And the two companies that published nothing did fine

The control group is instructive and slightly humiliating. Design Arena holds a paid Hugging Face "team" org with three members, seven followers and a tagline naming RL, alignment and human preference — and has published zero datasets, zero models, zero Spaces and zero papers. It has ~40 live leaderboards, ~146 GitHub stars led by agent-runner at 105, and a $7.9M seed led by Index. Taste Labs has no Hugging Face org, no GitHub org, no arXiv paper, no dataset, no benchmark and no code — six blog posts — and raised $18.5M co-led by CRV and Amplify.

Neither is evidence that publishing is worthless. Both are evidence that publishing is not the variable that separates a funded company from an unfunded one in this market, which is exactly the trap a plan built on "we will out-publish them" walks into.

Two things the record does not show

No model on Hugging Face declares training or fine-tuning on any Contra dataset, any Lica dataset, purvanshi/TASTE, or AfterQuery/ui-bench. FinanceQA has exactly one. Downstream adoption of design-preference data across this entire cohort is approximately zero.

No model card, system card or lab publication references any of it. That is negative search evidence and is [WEAK] as proof of absence, but it is consistent across the citation graph, the HF model index and the press record.

And the press record is its own finding: in this cohort, funding and M&A generate coverage and research does not. Design Arena got TechCrunch on a $7.9M seed; Lica got TechCrunch on its acquisition; AfterQuery got Forbes. Contra's 39 studies produced one launch-period newsletter, a directory listing and an automated ResearchGate mirror. On Hacker News, Contra has three submissions ever — the HCB study took 18 points and 2 comments — while Design Arena's two launch posts took 89 points / 29 comments and 74 / 24. Two posts beat thirty-nine studies.

The lesson, without hedging

Volume bought Contra nothing. Not citations, not downloads, not downstream training use, not press, not Hacker News, not a single referenced model card. It plausibly bought inbound sales conversations that are invisible from outside — and Contra's own job specs say that is the point, with the Strategic Project Lead owning "case studies and benchmark publications" as a demand-generation line. Read as sales collateral with a research budget, the programme may well be working. Read as research, it did not travel.

One rigorous artefact beats thirty-nine studies

The evidence says the unit of account is not the study. It is the artefact that someone else can run: a dataset with a licence, a paper with a method, a repository with a harness, and a number that survives being re-derived. FinanceQA is 148 rows with a repo and 17 citations. Contra is 39 studies, 8 datasets, 2 papers, no repo and zero. Build one thing properly.

That is the whole argument for the sequencing in Rebuild the Human Creativity Benchmark: take something that already exists, do it properly, ship the code with it. The full record is at Every public artefact, by organisation, the mechanics of making a release travel at How an artefact travels, and what nobody has built at What nobody has built. The strategic frame these numbers sit inside is The taste read and What to build first; the reason a benchmark is customer acquisition rather than product is Publishing the benchmark.