The competitor that decides whether this business exists is not another expert-data startup. It is the layer that collects the same raw material as a free by-product of a product people already want to use.
6,047,075 image votes and 645,773 video votes on Arena's public leaderboards as of 25 August 2026 (arena.ai; arena.ai) — 76 image models, 46 video models. $100M annualised revenue on 28 people (TechCrunch, 29 Jun 2026; Contrary). Marginal cost per vote: server time.
Set the first number against the specialist output it competes with. Contra Labs' flagship benchmark is 5,940 pairwise judgements from roughly 30 evaluators (arXiv 2606.30561); its Design Crit / TASTE set is 14,400 comparisons from 10 designers (arXiv 2605.20731). Arena's image board alone holds roughly 420× Contra's entire published output, and Arena paid nothing for it while Contra paid designers about $90/hr.
What Arena is
Arena Intelligence Inc., incorporated April 2025 out of the 2023 Berkeley Chatbot Arena project, San Francisco, founded by Anastasios Angelopoulos, Wei-Lin Chiang and Ion Stoica (Sacra; Contrary). It raised a $100M seed in May 2025 (a16z, UC Investments) at $600M post and a $150M Series A on 6 January 2026 (Felicis, UC Investments) at $1.7B post — $250M in about seven months (TechCrunch).
Traffic went 1M monthly visitors (April 2025) → 5M (January 2026) → 10M (June 2026), with 82M cumulative votes and 700M cumulative conversations (Sacra).
Revenue went $30M annualised in December 2025 — under four months after the first paid product — to $100M annualised in June 2026. The paid product is AI Evaluations: private pre-release arenas, custom eval tooling and analytics, API/SDK access, and, the line that matters here, "curated datasets for reward model training" (Contrary). Pricing is consumption-based on "volume of battles, votes, and prompts consumed" — so it is project revenue with a re-run cadence, not subscription, and the CEO says so (Sacra; TechCrunch). Apply GMV is not revenue before repeating any of it as ARR.
And they are not giving the corpus away. Contrary, verbatim: "LMArena does not monetize its entire dataset; it commits to publicly releasing up to 20% of preference data to support open research, while retaining the remaining 80% as a strategic asset" (Contrary).
That sentence is the competitive position in full. Arena is a consumer product that generates a proprietary preference corpus at zero marginal cost, sells access to a fifth of it as goodwill, sells reward-model training data from the rest, and has 28 people.
The volume fight is over and the specialist lost it
There is no version of a professional panel that reaches these numbers. Using the one disclosed cost anchor in the field — designers paid "a flat project fee that averaged approximately $90 per hour" producing 14,400 comparisons in about 145 designer-hours, i.e. ~$0.91 per comparison (arXiv 2605.20731) — Arena's image board would have cost a specialist about $5.5M in direct annotator labour.
| Corpus | Comparisons | Cost at $0.91 each | Actually paid |
|---|---|---|---|
| Contra HCB | 5,940 | ~$5,400 | ~$5,400 |
| Contra TASTE / Design Crit | 14,400 | ~$13,000 | ~$13,000 |
| HPD v2 | 798,090 | ~$726,000 | undisclosed, 57 contractors |
| HPDv3 | 1,170,000 | ~$1.06M | undisclosed, professional artists |
| Arena text-to-image | 6,047,075 | ~$5.5M | ~$0 |
Two things follow. First, a frontier-competitive professional preference dataset costs roughly $1M of direct annotator labour — a fundable line item, not a moat, and a build-versus-buy decision a lab makes in one meeting. Second, competing on comparison volume against a free consumer funnel is a losing trade at any margin.
What is actually left, argued properly
The temptation is to dismiss the arenas as noisy amateurs. That is wrong, and a buyer will know it is wrong. Arena's data is genuinely better than a small panel's on several axes: it is larger by orders of magnitude, it is continuously refreshed, it covers 150 countries, it is demographically broad, and it measures the thing most consumer models are actually optimising — average user delight. Design Arena adds geographic and temporal drift on top. If a lab's target metric is aggregate user preference, buying panels instead of arenas is malpractice.
The specialist survives only where a click cannot carry the information. Four places, in descending order of defensibility:
1. Rater identity. An arena vote is anonymous and uncredentialed by construction. A rating attributable to a named art director with a decade of client work is a different evidentiary object, and it is the only one that survives a question like "whose standard is this?" — the question that decides whether a lab can defend a model card claim. This is why Telling a good designer from a confident one is a product page and not an ops page.
2. Written rationale. Arenas capture a preference; they do not capture a reason. A lab debugging why its model's typography reads as cheap needs sentences, not Elo. Contra's flagship produced 3,675 written rationales alongside its 5,940 judgements (arXiv 2606.30561) — the rationales, not the judgements, are the part Arena cannot generate.
3. Multi-axis rubrics. A single winner collapses prompt adherence, craft, originality and usability into one bit. Nine-dimension scoring separates them, which is what turns a leaderboard position into an engineering instruction. See Phase decomposition for the version of this that actually sells: the finding that model rankings invert between ideation, mockup and refinement.
4. Captured process. No arena collects trajectories — screen recordings, edit sequences, narrated reasoning. This is the only asset in the category with a genuine cost floor, because it requires a standing relationship with professionals doing real work. The trajectory moat and Eight units and one cost anchor carry it; What to build first argues it is the whole wedge.
The arena layer wins on volume permanently. Everything the specialist sells must be something that cannot be expressed as a click: a name, a reason, an axis, or a recording. Any product line that reduces to "we also have preference pairs" is priced against $0 and will lose.
Yupp, the corpse that proves both halves
Yupp.ai raised $33M seed from a16z crypto (Chris Dixon), founded June 2024, and wound down on 15 April 2026 (The AI Cemetery; AIny). Users compared outputs from hundreds of models, were paid in reward points redeemable through payment rails, and Yupp aggregated the rankings and sold them to labs. It reached ~1.3M signups and "millions of preferences monthly".
The stated reason for death is the single most useful sentence in this document: the market "shifted from chatbot comparison toward expert feedback, domain-specific labeling, and agentic systems" (The AI Cemetery; Valasys).
Read it twice, because it cuts both ways.
For the thesis: generic crowd preference stopped being purchasable. The exact shift Yupp died of — toward expert, domain-specific feedback — is the client's business case, validated by a company that failed for the opposite reason.
Against it: Yupp had 1.3M users, millions of monthly preferences, $33M of a16z money and a payment rail, and still could not find buyers at a sustaining price. The demand side for preference data is thinner and far more concentrated than the supply-side story implies. AI Weekly makes the same connection when assessing Design Arena's revenue claim (AI Weekly).
The honest synthesis is that this market has one large winner, one funded challenger, one corpse, and eighteen months of history. That is not a market with room for many vendors selling the same undifferentiated thing. How much money is actually in the buyer pool and Sizing the taste market carry the sizing.
Artificial Analysis: a near-total blank
The third name in this layer needs stating honestly rather than assessed.
Artificial Analysis runs an Image Arena and a Video Arena with a published methodology: "users compare two videos generated from the same prompt by different models and select the one they prefer", aggregated by Bradley-Terry MLE, rescaled to an Elo-like range, recomputed hourly (artificialanalysis.ai). Co-founders Micah Hill-Smith (CEO) and George Cameron (CPO); Crunchbase lists Newark, Delaware, headcount 11–50 and about 2.4M monthly web visits (Crunchbase).
No disclosed funding on the Crunchbase profile. No revenue, no business model, no vote counts, no voter qualifications, no voter compensation, and no data-licensing or data-sales statement anywhere in the methodology (artificialanalysis.ai). The reasonable working assumption is that Artificial Analysis is a benchmarking publication with a reputation asset rather than a preference-data vendor — but that is an inference from silence, and the absence of a data-sales statement is not evidence that none exists. This is the largest single hole in the competitive map of this layer.
Which leaves the layer looking like this: one company with $250M, 82M votes and an explicit reward-model data product; one seed-stage challenger with 5.3M users and an unverifiable revenue claim; one publication whose commercial posture nobody outside it knows; and one grave.