The artefact is the easy part. Everything on this page is about whether anyone ever runs it.
Licence
Pick one, apply it everywhere, and never ship without it.
Contra uses CC-BY-4.0 across all eight datasets. That is the right default for a marketing artefact: permissive, familiar, attribution-preserving, and it means a lab's counsel clears it in a morning. Lica mixes Apache-2.0 on the two GDB repos with CC-BY-4.0 on lica-dataset and MIT on purvanshi/TASTE, which is untidy but harmless.
AfterQuery is the cautionary case: three of six datasets carry no licence at all — ui-bench, App-Bench and MCP-Universe. For a company at a $3.2B valuation selling data governance as part of its pitch, that is a real defect and worth naming rather than glossing. An unlicensed public dataset is not permissively licensed; it is a dataset with no grant attached, which means a careful buyer's counsel says no and a careless one creates an exposure. UI-Bench is the most rigorous public human-judged design benchmark in the market and it is the one with nothing in the licence field.
The choice that matters for a plan built on a rebuilt benchmark: CC-BY-4.0 on the dataset, a permissive OSI licence on the code, and a stated licence on the model weights. Three fields, filled in on day one. The most common failure in this record is not choosing badly; it is leaving the field empty.
Ship the three things together
Every cited artefact in this inventory has a repository beside it. Every uncited one does not. AfterQuery's FinanceQA (17 citations) ships with FinanceQA; VADER (8) ships with vader; UI-Bench (2) ships an open framework. Contra Labs has no GitHub organisation at all — /orgs/contralabs and /orgs/contra-hq both 404 — and holds zero citations on two papers. For a company whose product is evaluation, that means there is no way for a lab to run a Contra evaluation.
Dataset, paper and code are one release, not three. A dataset without code is a file someone has to reverse-engineer. A paper without a dataset is a claim. Code without either is a tool with no demonstration. What gets cited is the combination, because a citation usually happens when someone ran the thing.
The bundle for a benchmark release, in the order a reader encounters it: a live page with the leaderboard and dates on it; the paper on arXiv; the dataset on Hugging Face with a datasheet naming panel composition, recruitment channel, pay and screening threshold; the repository with the harness, the scoring code and the reliability computations; the model weights with a training script; and a described-but-sealed test split with a submission endpoint. Six artefacts, one day, one announcement.
Add the small mechanical things that cost nothing. Point the paper's data link at the repo you will actually keep — HCB's paper cites contra-labs/HCB, an org that does not exist, saved only by a 307 redirect. Do not label the flagship a "preview release" — seven of Contra's eight cards do, which tells a reader the thing is not finished and there is no reason to depend on it. And keep modifying the repo: LMArena's leaderboard-dataset was last modified on the day the inventory was pulled and has 126,168 downloads; every Contra dataset was last touched within five days of creation and none since 22 July 2026. A repo that has not moved in six weeks reads as abandoned, and in this record the reading is roughly correct.
What the like-to-download ratio actually tells you
The intuition is that likes without downloads mean attention without adoption — people are liking a repo off a social post and nobody is building on it. It is a reasonable diagnostic and it is worth watching. It is also not what this record shows, and the honest version is more useful than the tidy one.
| Repo | Downloads (all-time) | Likes | Ratio |
|---|---|---|---|
contralabs/descript-video-editing-trajectories | 380 | 0 | — |
contralabs/premiere-video-editing-trajectories | 1,415 | 12 | 118:1 |
| Contra Labs, whole org | 5,133 | 24 | 214:1 |
| Lica World, whole org | 8,905 | 6 | 1,484:1 |
| AfterQuery, whole org | 14,545 | 31 | 469:1 |
lmarena-ai/arena-human-preference-55k | 40,661 | 159 | 256:1 |
Across Contra's org, likes run 0.004 to 0.014 per all-time download. descript-video-editing-trajectories is the largest trajectory repo by volume — 803 steps, 389 MB — and has zero likes, while premiere at a third the size has twelve and three times the downloads. The high-likes, low-downloads signature appears nowhere in this market. What appears instead is uniformly low absolute numbers for the design specialists and uniformly high ones for the arena, and the ratios are indistinguishable between the cited corpus and the uncited one.
So use the ratio as a watch metric, not a decision metric, and use the diagnostic that does discriminate: downloads with no declared downstream use. Zero models on Hugging Face declare training or fine-tuning on any Contra dataset, any Lica dataset, purvanshi/TASTE or AfterQuery/ui-bench. FinanceQA has exactly one. That number — declared dependents in the HF model index — is the adoption metric this market is actually failing, and it is the one to instrument from the first release.
Where the leaderboard lives
A leaderboard in a PDF is a table. A leaderboard on a page with dates on it is an instrument people return to.
AfterQuery hosts each benchmark on its own domain — uibench.ai, marketbench.ai, ide-bench.com, appbench.ai — with live boards and confidence intervals. Design Arena runs ~40 continuously-updated boards and is the most-discussed company in this cohort on Hacker News despite publishing no data at all. Contra has no leaderboard: the public HCB page describes a five-axis rubric that does not match the three-axis rubric in its own paper, publishes no prompt count, no model count and no update cadence, and "Creative Arena" is a marketing name for the research index — contra.com/creative-arena redirects to the research page.
The scale reference for hosting is LMArena: its arena-leaderboard Space has 4,980 likes, two orders of magnitude more than any other artefact in this document, and it sits on ~47,000 GitHub stars. That is not a target. It is the demonstration that the hosted, live, continuously-updated board is the artefact people attach to, and that the attachment accrues to the host rather than to the data.
Practical choice: a Hugging Face Space is free, discoverable inside the ecosystem where the dataset already lives, and likeable — which is the one place a like is a real signal, because it means someone bookmarked a tool. Note that Contra, Lica, AfterQuery and Design Arena between them operate zero Spaces. The field is empty.
Cadence
Contra published 39 dated studies in 140 days — one every 3.6 days, monthly at 6, 12, 3, 11, 5. It is an impressive operational achievement and it produced roughly 1.3 dataset downloads per study, zero citations and three Hacker News submissions ever.
The reason it compounded into nothing is structural rather than a matter of effort. No two studies used the same instrument. Different panels, different prompts, different rubrics each time, so the 39 results cannot be stacked into a trend line, a meta-analysis or a single claim. Thirty-nine incompatible snapshots are thirty-nine blog posts.
The alternative cadence is AfterQuery's: one domain benchmark roughly every four months for eighteen months, each naming specific commercial tools and showing them failing, each with a live board, each runnable by a stranger. Six papers, 32 citations, and a company that went from YC W25 to $100M annualised over the same window. Causation is not established and should not be claimed; the contrast in what the two publishing programmes left behind is not in dispute.
So: publish less, and publish the same instrument repeatedly. A quarterly re-run of one frozen benchmark against each new model slate is a series — it gets more valuable with age, it is the natural shape of a subscription, and it is the longitudinal artefact nobody in this market has. Filling the calendar with studies is a demand-generation programme, and Contra's own job specs say that is exactly what it is: the Strategic Project Lead owns "case studies and benchmark publications". Judged as sales collateral it may be working. Judged as research it left nothing behind, which is the finding at What publishing actually bought them.
The release this all applies to is at Rebuild the Human Creativity Benchmark; the record it is derived from is at Every public artefact, by organisation; and the reason a benchmark is customer acquisition rather than product is at Publishing the benchmark and What to build first.