Miju Labs

The security dossier

Every public artefact, by organisation

The full record for six organisations: every dataset with repo id, date, licence, size, downloads and likes; every paper with arXiv id and citation count; every repository with stars. Design Arena holds a paid Hugging Face org that is completely empty, Taste Labs has published nothing at all, and Gamma's org is empty while the Lica repos it bought are still live.

high confidence10 minupdated 2026-08-30inventory · datasets · papers · repositories · hugging face · arxiv

Everything public, as pulled from the Hugging Face datasets API (expand[]=downloadsAllTime), the dataset-viewer size API, the arXiv API, the Semantic Scholar Graph API and GitHub org listings on 2 September 2026. Download and like figures are snapshots and will drift. mo/all = last-month / all-time downloads.

Read this beside What publishing actually bought them, which is what these numbers mean, and What nobody has built, which is what is missing from them.

Contra Labs

Eight datasets, all CC-BY-4.0, all created between 18 May and 17 July 2026, all last modified within five days of creation and none touched since 22 July 2026. Every card except HCB is labelled a "preview release". Every card carries UTM-tagged links back to contralabs.com and a "Working with Contra Labs" section with a partnerships address — these are lead-generation assets with a dataset attached, and they say so.

ArtefactCreated / modifiedLicenceSizemo/all · likes
contralabs/HumanCreativityBenchmark18 May / 24 Jun 2026CC-BY-4.08,012 rows over 5 CSVs — 95 prompts, 380 outputs, 3,174 pairwise, 2,116 scalar, 2,247 qualitative278 / 532 · 2
contralabs/photoshop-creative-design-trajectories2 Jun / 10 Jun 2026CC-BY-4.0294 steps, 163 MB232 / 822 · 3
contralabs/firefly-creative-campaign-trajectories9 Jun / 10 Jun 2026CC-BY-4.0137 steps, 63 MB101 / 399 · 1
contralabs/gemini-creative-campaign-trajectories9 Jun / 10 Jun 2026CC-BY-4.0266 steps, 116 MB150 / 432 · 1
contralabs/video-detail-annotation1 Jul / 2 Jul 2026CC-BY-4.015 rows — 3 video models, timestamped editor comments369 / 1,007 · 3
contralabs/premiere-video-editing-trajectories14 Jul / 16 Jul 2026CC-BY-4.0234 steps over 4 trajectories, 182 MB614 / 1,415 · 12
contralabs/creative-ad-design-dataset17 Jul / 22 Jul 2026CC-BY-4.035 rows — brand brief → approved ad, 126 MB101 / 146 · 2
contralabs/descript-video-editing-trajectories17 Jul / 18 Jul 2026CC-BY-4.0803 steps, 389 MB283 / 380 · 0
The Human Creativity Benchmark, 2606.3056129 Jun 2026, 30 pparXiv default28 evaluators, 93 prompts, 80 sessions, ~15.5k judgments claimed0 citations
TASTE, 2605.20731 (with Lica)v1 20 May, v2 2 Jun 2026CC-BY-4.010 designers, 9 criteria, 80 prompts each, 21,600 pairwise0 citations
contralabs.com/research8 Apr – 26 Aug 2026none stated39 dated studies + 5 dataset pages + 1 methodology pagen/a
HF org contralabs8 datasets, 0 models, 0 Spaces, 0 papers, 5 followers
GitHubNo organisation exists/orgs/contralabs and /orgs/contra-hq both 404

Org totals: 5,133 all-time downloads, 24 likes. Total trajectory volume across the five trajectory repos: 1,769 steps.

Two details worth carrying. The trajectory schema is the strongest thing in the whole inventory — per-step image, thought transcribed from the editor's spoken narration rather than model-synthesised, action_type, structured tool_call, execution_paths (MCP tool / keyboard shortcut / menu path) and a preferred_execution marker (The trajectory moat). And the HCB paper's data link points at huggingface.co/datasets/contra-labs/HCB, which 307-redirects to the live repo — the link works, but the org contra-labs does not exist, so the paper cites a repo name Contra abandoned.

Lica World, now Gamma

ArtefactDateLicenceSizemo/all · likes
lica-world/GDB13 Apr / 5 May 2026Apache-2.033,886 rows, 17.0 GB, 40 task configs439 / 4,615 · 1
lica-world/lica-dataset14 Apr 2026CC-BY-4.01,148 compositions — a 0.07% sample of the 1,550,244 in the paper111 / 2,710 · 4
lica-world/GDB-video1 Jul 2026Apache-2.0654 rows over 136 animated layouts; reference/sora2/veo31 configs393 / 1,580 · 1
purvanshi/TASTE28 / 29 May 2026MIT36,695 rows — 14,460 rankings, 3,200 hallucination flags, 721 prompts, 644 assets, 10 evaluators408 / 1,809 · 9
LICA, 2603.1609817 Mar 20261,550,244 compositions, 971,850 templates, 27,261 animated5 citations, 4 self
GDB, 2604.041925 Apr 202650 tasks over 5 axes (site says 49)4 citations, 3 self
SVG structural metrics, 2604.088099 Apr 202619,000+ edits, 5 edit types, 5 systems1 citation (self)
Evaluating Design Video Generation, 2605.1622315 May 2026 · ICML 2026 Workshop on Human-AI Co-Creativity4 evaluation dimensions0 citations
A Typography Benchmark for Co-Creative Graphic-Design Agentsforthcomingno arXiv id; visible only as a citing edge in Semantic Scholar
purvanshi/lica-benchLICENSE present, type not surfaced45 tasks / 39 benchmarks / 7 domains over 1,148 LICA layouts10★, 2 forks
purvanshi-lica/tasteMITanalysis/ + data/ + taste-scorer (pip-installable, Qwen3-VL-Embedding-2B backbone)3★, 1 fork

Note that TASTE and both repositories live on personal accounts, not the org — which matters now that Lica belongs to Gamma. taste-scorer ships code and data and no trained checkpoint.

AfterQuery

Six datasets, 23 followers, seven org members — the largest Hugging Face footprint of the specialist-data companies, and the only corpus in this document with real citation counts.

ArtefactDateLicenceSizemo/all · likesCitations
AfterQuery/FinanceQA29 Jan / 21 Feb 2025Apache-2.0148 rows187 / 6,311 · 2117 (1 influential)
AfterQuery/vader16 / 23 May 2025CC-BY-4.02,659 files; 174 real vulnerabilities286 / 5,040 · 48
AfterQuery/ui-bench25 Aug 2025none set30 client-style briefs, 5 categories44 / 395 · 22
AfterQuery/App-Bench10 Dec 2025none set6 rows98 / 627 · 4
AfterQuery/MCP-Universe29 Jan 2026none set2 rows — spreadsheet tasks from Goldman Sachs / Evercore experts54 / 549 · 0
AfterQuery/OmicsBench-grader-databases10 Apr / 3 May 2026MITsynonym databases for deterministic biology grading46 / 1,623 · 0
Market-Bench, 2512.1226413 Dec 20253 strategies, 13 models, 5 rounds1
IDE-Bench, 2601.2088628 Jan 202680 tasks, 8 never-published repos2
SpreadsheetBench 2, 2606.2995529 Jun 2026321 tasks, best model 34.89%2 (1 influential)
GitHub org, 11 reposmixedharbor (2★, Apache-2.0, official harness for Terminal-Bench 2.0), vader (11★), FinanceQA (9★), anvil (8★), IDE-Bench (5★), swift-anvil (2★), appbench.ai-docs (1★)~38★ total

MCP-Universe itself is Salesforce AI Research's benchmark (arXiv 2508.14704), not AfterQuery's — the HF repo is a two-row contribution to it. And three of six datasets carry no licence at all, which for a $3.2B-valuation company is a governance gap and a defect worth naming (How an artefact travels). harbor has 1,286 commits and two stars: heavily built, barely adopted, and it contains no creative or design environments.

Design Arena

ArtefactDetail
HF org designarenaexists on a paid "team" plan, 3 members, 7 followers, tagline "RL, fine-turning, model alignment, human preference" — and 0 datasets, 0 models, 0 Spaces, 0 papers
GitHub org, 7 reposagent-runner (105★, MIT), html-to-pptx (21★, MIT), MicroEvals (7★), audio-arena-bench (6★), LeWitt-Bench (4★, a Sol LeWitt instruction-art dataset), audio-agent-bench (3★), docs (0★) — ~146★ total
Leaderboardsdesignarena.ai, ~40+ boards, Elo / Bradley–Terry, continuously updated from public votes. No vote total, no downloadable eval set, no public dataset
notes.designarena.ai12 posts, 23 Feb – 4 Aug 2026, including Fullstack Arena Methodology and the Audio Realism Benchmark
arXivnone, under any of Design Arena / Intelligence / Arcada Labs

A paid team org that is completely empty is the single most legible negative in the inventory. It is consistent with two readings — a distribution plan provisioned and not yet executed, or preference data being distributed privately to buyers — and the record does not separate them (What this plan could not establish). Note also that Design Arena has open-sourced tools and never open-sourced preference data: html-to-pptx is a HTML-to-PowerPoint conversion library, not a corpus. ArcadaLabs on Hugging Face is an unrelated 2022 account hosting two MLX conversions and is not Design Arena's parent; earlier notes identifying it as such should be corrected. Full read at Design Arena, in full.

Taste Labs

ArtefactDetail
Hugging FaceNo org. tastelabs, taste-labs and TasteLabs all 404
GitHubNo organisation found
arXivNo paper
Datasets, benchmarks, leaderboards, codeNone
BlogSix posts, 16 June – 16 August 2026
Requests for research16 Aug 2026, 14-min read, Hamidah Oderinwale — see The eleven problems a rival published
Prototype fellowship16 Aug 2026 — grants, compute, custom human data, collaborative publishing

Taste Labs has published no technical artefact of any kind and raised $18.5M co-led by CRV and Amplify. It is nonetheless the more dangerous competitor for reasons that have nothing to do with this table (Taste Labs, in full).

Gamma

gamma-app is a verified Hugging Face org with 3 members, 6 followers and 0 datasets, 0 models, 0 Spaces, 0 papers. Gamma acquired Lica on 25 August 2026 (TechCrunch); lica.world remains live with five blog posts and the GDB explorer, and all three Lica HF repos plus both personal-account repositories are still public. Nothing has been migrated, withdrawn or relicensed. That is a window, not a policy, and it may close.

LMArena, as a ruler

ArtefactDateSizemo/all · likes
lmarena-ai/VisionArena-Chat14 Jan / 4 Feb 2025230K VLM conversations3,504 / 162,688 · 15
lmarena-ai/leaderboard-dataset2 Apr 2026, last modified 2 Sep 20262,277,369 rows, 21 configs incl. text_to_image, text_to_video, image_edit, webdev25,144 / 126,168 · 23
lmarena-ai/search-arena-24k15 May 2025 / 3 Mar 202624k465 / 81,199 · 40
lmarena-ai/webdev-arena-preference-10k15 Jan 202510k webdev preferences348 / 78,838 · 20
lmarena-ai/arena-human-preference-140k1 Aug 2025140k2,568 / 56,505 · 55
lmarena-ai/arena-human-preference-55k2 May 202455k1,867 / 40,661 · 159
lmarena-ai/arena-human-preference-100k4 Feb 2025100k811 / 26,713 · 49
lmarena-ai/Arena-T2I-Hard30 Jun 2026310 prompts, ~30 decomposed yes/no constraints each84 / 268 · 1
Org totals2023 – present27 datasets, 75 models, 10 Spaces across lmarena-ai, lmarena, lmsysarena-leaderboard Space: 4,980 likes
GitHubFastChat (40k★), RouteLLM (5.4k★), arena-hard-auto (1.1k★), copilot-arena (364★), p2l (275★), arena-rank (109★), PPE (65★), search-arena (58★)~47,000★

Arena-T2I-Hard is the strategically interesting one: its own paper finds that high public-arena rankings fail to predict faithfulness — LMArena publishing the critique of its own headline metric and shipping the replacement. See The arena layer.

The reference point that closed a gap

PatronusAI/figmatrace plus Qwen3.8-27B-Figmatrace-SFT, created 19 August and last modified 25 August 2026, CC-BY-4.0: 3,469 design trajectories, 22.3 GB, 200+ hours of expert video, 126 tasks, at 1,784 downloads and 12 likes in a fortnight; the model at 19 downloads (arXiv 2608.21460). It is the largest open design-trajectory corpus in existence and it did not come from anyone in this inventory.

What this inventory does not contain

No pricing. Not one organisation here publishes a rate for a preference dataset, a trajectory hour or an evaluation engagement. No downstream usage. Zero models on Hugging Face declare training on any Contra dataset, any Lica dataset, purvanshi/TASTE or AfterQuery/ui-bench; FinanceQA has exactly one. And no aesthetic reward-model weights from anybody — the most conspicuous absence in the whole record, treated at What nobody has built.