Contra Labs
Not a startup — a business line of a six-year-old freelance marketplace, launched five months after a $740K cheque. The template everyone wants to copy, and the parts of it that do not survive inspection.
Latest (Sep 2026): Contra launched Contra Labs on 31 Mar 2026 as 'the first frontier data and evaluation lab for creative AI'. It draws on a claimed 1.5–1.7M-member creative network and offers a Creative Arena, the Human Creativity Benchmark, preference data and screen-recorded trajectories. Research followed in June: Design Crit (17 Jun; off-the-shelf VLM judges below 55% agreement with designers, a small trained head at 61.1%), the TASTE dataset co-built with Lica World, and the HCB paper on arXiv (29 Jun). Its modelling partner Lica was acquired by Gamma on 25 Aug 2026. In Aug 2026 it was hiring Research Scientists (Human Data & Evaluation) and project leads at $150–200K in NYC and SF. No frontier-lab customer has been named publicly.
Contra Labs is not a startup. It is a business line of Contra.Work Inc., a Delaware corporation incorporated in 2018 that runs a commission-free freelance marketplace (SEC Form D). The homepage says so in the footer — "Powered by Contra" — the LinkedIn link points at linkedin.com/company/contrahq, and the jobs board is jobs.ashbyhq.com/contra with Contra Labs appearing as a department inside it (contralabs.com; Ashby posting API). EDGAR full-text search returns zero hits for "Contra Labs" across all form types (EDGAR FTS). There is no separate legal entity.
That correction reframes what is actually being copied. The launch was 31 March / 1 April 2026 (Contra blog; Creator's Toolbox, 1 Apr 2026) — five months after a $740,000 venture cheque from Zentavo VC in October 2025 (Clay funding dossier). Before that: a $14.5M Series A led by Unusual Ventures in February 2021 and a $30M Series B led by NEA in November 2021, with Cowboy Ventures participating (TechCrunch, Feb 2021; TechCrunch, Nov 2021). Total raised is about $45.2M (Clay).
A company that raised $30M in 2021 and takes $740K four years later is not on an up-and-to-the-right path. The funding shape says the marketplace stalled and this is the second act — an incumbent repointing existing supply, not a new company assembling one. Every "how did they get 1.5M designers so fast" question has the same answer: they did not. It took six years and $45M, paid for by a different business.
The Crunchbase profile at crunchbase.com/organization/contra-labs is a different, defunct company: a Brazilian app developer founded in 2007 in Salvador, Bahia by Victor Cardozo and Gustavo Carvalho, listed as permanently closed, whose contact address was contato@contralabs.com (Crunchbase). Contra evidently acquired the domain from them. The correct parent profile is Crunchbase — Contra.Work Inc. Anyone building a competitive map from Crunchbase alone will import a dead Brazilian company's founding date into their model of the Design and UI/UX vertical.
What they actually sell
Three commercial lines and one marketing line, and the ordering on the homepage is backwards.
Preference pairs and rubric scores are the headline product: "Seed datasets (image, video, UI, motion). Corrective SFT data and RLHF preference pairs" (contralabs.com). The flagship study alone produced 5,940 pairwise judgements and 3,675 written rationales, plus 1–5 Likert scores on Prompt Adherence, Usability and Visual Appeal (arXiv 2606.30561). The public rubric is five-axis — visual quality, prompt adherence, originality, utility, and motion realism for video — with three or more professional evaluators per output (contralabs.com/human-creativity-benchmark).
Design Crit is the reward-model line. "Criteria-Resolved Image Taste" is a designer-annotated preference set scoring nine dimensions rather than one verdict, on which they trained "a small pairwise-difference head on top of a frozen vision-language encoder, with no fine-tuning of the backbone" — their words, a "deliberately modest model" (Design Crit). It is explicitly "a Lica × Contra collaboration": Lica World is a separate creative-AI research company positioning itself as "built for research labs developing the next generation of design AI" (lica.world/blog). The modelling capability was partnered in, not built.
Screen-recorded trajectories are the differentiated asset, and they are buried third. The Premiere Pro dataset card gives the shape: 234 steps across 4 trajectories, 111–245 minutes per session, recorded on macOS while professional editors built vertical social reels from real client briefs. Per-step schema: trajectory_uuid, session_uuid, image, thought, action_type, tool_call as structured JSON, execution_paths (MCP tool / keyboard shortcut / menu path) and preferred_execution (HF card). The card's claim is the whole pitch: "The screenshots, the recorded action, and the pointer coordinates are the editor's real execution" — and the thought field comes from the editor's spoken narration, explicitly distinguished from "model-synthesized rationales".
That last detail is the one to take. You cannot synthesise a professional's mouse path through Photoshop, and you cannot obtain their reasoning without them talking while they work. Preference pairs are a commodity — any lab with thirty contractors and two weeks makes them. The narrated trajectory requires a standing relationship with professionals doing real client work, which is exactly what Which side you build first says money cannot buy quickly. Contra owns the moat and leads with the commodity.
gemini-creative-campaign-trajectories and firefly-creative-campaign-trajectories are named for the tool the designer used, not the customer (HF org). They are not Google or Adobe deals.
The phase decomposition, which is the smartest thing here
The Human Creativity Benchmark splits creative work into three phases — Ideation → Mockup → Refinement — and shows that model rankings invert between them. Claude leads Ideation on landing pages, Gemini dominates Mockup at a 68.9% win rate, Claude reclaims Refinement (HCB). Veo 3.1's 61% ideation win rate falls to 39% at refinement (contralabs.com/research).
The methodological claim underneath is that evaluator disagreement in creative domains is signal, not noise — it separates convergence on shared professional standards (typography, layout, hierarchy) from divergence in legitimate taste (arXiv). Agreement is highest on Prompt Adherence and lowest on Visual Appeal, as you would expect if the claim holds.
This matters commercially, not intellectually. A leaderboard tells a lab "you are third", which is useless and unwelcome. A phase decomposition tells them "your model collapses at refinement, here is why", which is a purchase order. It is portable: every domain has an analogous split — issue-spotting → drafting → redlining in Law, differential → workup → management in Clinical medicine. Finding yours is the transferable move; see The specialist wedge.
The buyer gap
The job posts describe the customer as "frontier AI labs and product companies" and "frontier AI research teams" (Ashby; Ashby). No frontier lab is named as a customer anywhere — not OpenAI, not Anthropic, not Google, not Meta, not xAI. There is no logo wall, no case study page, no testimonial.
The only corroborated partner list comes from a launch newsletter, verbatim: "Contra Labs is already partnering with a wide range of AI-powered creative tools (including Framer, Webflow, Lovable, Replit, HeyGen, and others)" (Creator's Toolbox). [UNVERIFIED] — it could not be corroborated from Contra's own site, and "partnering with" spans everything from a signed contract to a shared Slack channel.
Every name on that list is an application-layer AI company, not a model lab — Series B/C startups with product budgets in the tens of thousands, not labs with nine-figure data budgets. Mercor's revenue comes from OpenAI, Google DeepMind and Meta (Sacra). If Contra Labs' real buyer base is design-adjacent app companies, the addressable spend is one to two orders of magnitude below what the "Mercor for design" framing implies. That is the most important risk here, and the thing to stress-test first in your own niche — see How much money is actually in the buyer pool.
The 39 published studies evaluate Claude, Gemini, GPT, Qwen, Veo, Seedream, FLUX, Kimi, Grok, Nano Banana and Meta Muse — as subjects, not as clients (contralabs.com/research). "Where four AI models break when they build a landing page" is a lead-generation tactic aimed at those labs, not evidence of a relationship with them.
The supply terms — the IP is the weapon, not the rate
| Term | What Contra Labs offers | Source |
|---|---|---|
| Rate | Up to $100/hr, depending on role and seniority; some briefs flat-fee | contralabs.com/jobs |
| Brief length | Typically 5–10 hours, occasional longer engagements | Help Center |
| Payment | Within 7 business days of task completion | Help Center |
| Platform fee | None taken from the expert's rate | contralabs.com/jobs/photographers |
| IP | Expert licenses work to Contra Labs — not an assignment | contralabs.com/jobs/photographers |
| Credit | Published work carries an author byline | contralabs.com/jobs/photographers |
| Refusal | Experts "can decline projects that don't feel appropriate" | contralabs.com/jobs/photographers |
| Exclusivity | None found in any recruiting page, help article or job post [WEAK on the negative] | — |
The rate is not the recruiting weapon. Up to $100/hr sits above Mercor's $85+/hr average and well above micro1's $30–65/hr creative band (Sacra; aitraining.jobs) — but it is below what a good brand designer bills a real client. On price alone, Contra Labs is fill-in work.
The IP terms are the weapon. Licence-not-assignment, byline credit, no fee skimmed off the rate, the right to decline, seven-day pay instead of thirty to sixty. In a profession ideologically wary of training the thing that replaces it — Envato's survey of 1,780 creatives found graphic designers and illustrators have the lowest daily AI adoption at 40% and the highest frustration (Envato) — those terms are cheaper than a 30% rate premium and considerably more effective. The intake funnel matches: portfolio review, a video interview compressed in one role to a 30-second intro, and an assessment that is "evaluating AI-generated brand assets and providing structured feedback", about 10 minutes in total (Help Center). The work test is the work.
None of it buys lock-in. No exclusivity clause appears anywhere in the public funnel; experts keep a licence, can decline, and can take Mercor's rate tomorrow — the standard Getting cut out exposure of a marketplace that does not own the relationship.
The scale claim collapses on contact with the methodology
The homepage says "1.5M+ verified creative experts" in the hero and "1.7M+ creative experts" in the stats block, on the same page (contralabs.com). The benchmark page repeats the contradiction (contralabs.com/human-creativity-benchmark); the Creative Human Data page says 1.7M+ alongside "400+ skills and tools" and "$250M+ verified expert earnings" (contralabs.com/creative-human-data); the arXiv paper says 1.5M+ (arXiv).
Set either number against the studies. The flagship benchmark used 28–31 evaluators from 13 countries across 80 sessions, 93–95 prompts and 380 outputs from 13 models (arXiv; HF HCB card). Design Crit used 10 professional designers in two cohorts of five (Design Crit).
The 1.5M counts registered profiles on the parent marketplace, not the vetted Labs roster — nothing claims 1.5M people passed Labs vetting, and the vetted number is never disclosed. A technically literate buyer spots this in the first meeting. Worse, it advertises how cheap the studies are to replicate.
The org chart tells you the ceiling
| Role | Location | Band | Posted |
|---|---|---|---|
| Strategic Project Lead | NYC, onsite | $150K–$200K + equity | 2026-06-08 |
| Strategic Project Lead | San Francisco, remote | $150K–$200K + equity | 2026-07-01 |
| Special Project Lead | NYC, onsite | $150K–$200K + equity | 2026-07-30 |
| Research Scientist, Human Data & Evaluation | NYC, onsite | $150K–$200K + equity | 2026-08-17 |
| Research Scientist, Human Data & Evaluation | San Francisco, remote | $150K–$200K + equity | 2026-08-17 |
Source: Ashby posting API.
Five roles, all delivery or research-operations. Zero ML engineers, zero infrastructure engineers, zero applied scientists who train models. The band is flat at $150–200K including the PhD-preferred Research Scientist — a quarter to a half of what a frontier lab pays research staff. One posting requires "data labeling or annotation experience": they are hiring from Scale, Surge and Appen, not from design studios. Another owns "case studies and benchmark publications", which makes the research programme a demand-generation programme by its own job spec.
The interview loops name an org chart nobody is named in: Research Scientist runs recruiter → Head of Research → Data Lead → Head of Contra Labs → paid case study; Strategic Project Lead runs recruiter → CEO → technical → VP of Product → paid case study (Ashby; Ashby). The CEO personally interviews project leads — at five months old, Ben Huffman is running this himself.
Two reads. First, NYC Williamsburg, five days onsite, carved out of a parent that brands itself remote-first with no-meeting Tuesdays and Wednesdays (contra.com/careers) — the exemption a founder grants a unit that is the company's actual bet. Second, that org chart can only produce a consultancy. Delivery leads plus research-ops plus a rented model from Lica is an agency P&L with a data deliverable. If you want a reward model as a product, one of your first five hires has to train models.
Research as demand generation
39 studies in five months — roughly two a week — plus an arXiv paper and eight CC-BY-4.0 datasets on Hugging Face (contralabs.com/research; HF org).
| Dataset | Rows / steps | Downloads (all time) | Likes |
|---|---|---|---|
| HumanCreativityBenchmark | 8,012 rows | 532 | 2 |
| premiere-video-editing-trajectories | 234 steps (4 trajectories) | 1,415 | 12 |
| video-detail-annotation | 15 rows | 1,007 | 3 |
| photoshop-creative-design-trajectories | 294 steps | 822 | 3 |
| gemini-creative-campaign-trajectories | 266 steps | 432 | 1 |
| firefly-creative-campaign-trajectories | 137 steps | 399 | 1 |
| descript-video-editing-trajectories | 803 steps | 380 | 0 |
| creative-ad-design-dataset | 35 rows | 146 | 2 |
An earlier version of this page reported far larger figures in the downloads and likes columns. Those were the datasets' row and step counts read into the wrong columns, and the argument built on them — that likes wildly exceeded downloads, so this was attention rather than adoption — does not survive the real numbers. The corrected figures come from the Hugging Face API and are reproduced in Every public artefact, by organisation.
The eight repositories total roughly 5,100 downloads all time, against 24 likes. The flagship benchmark — the artefact the company's whole research programme points at — has 532 downloads and two likes. The best performer is a set of four Premiere trajectories. For comparison, Patronus AI's FigmaTrace shipped 3,469 design trajectories in August and took 1,784 downloads in a fortnight. See What publishing actually bought them for what that gap means.
The one sales statistic
Off-the-shelf VLM judges reach roughly 54% agreement with professional designers. Chance is 50%. The human single-rater ceiling is 74.1%. Contra's trained head reaches 61.1% (Design Crit).
That is the commercial argument in one line: your LLM-as-judge is barely better than a coin flip on design quality, we have the humans who fix it, and a deliberately modest model proves the data moves the number. It works because it measures a deficiency in the buyer's own stack rather than asserting something about Contra. Find your equivalent statistic before writing a line of outbound.
Pricing
Undisclosed. No pricing page, no tiers, no "starting at", no rate card anywhere on contralabs.com or contra.com. The only commercial front door is a Calendly link — calendly.com/partnerships-contra/contra-labs-partnership-request — the primary CTA on every page (contralabs.com; contralabs.com/creative-human-data).
A Calendly, rather than a contact-sales form or self-serve checkout, is the tell of a business doing a single-digit number of bespoke deals. With five open delivery roles and a five-month history, buyer One customer is a binary event is near-total by construction. Any margin figure quoted for this company is an inference from the cost side; see GMV is not revenue before repeating one.
Pricing to any buyer — no rate card, no tier, no observed contract value. Revenue, ARR, bookings or gross margin — nothing public at either Contra Labs or the parent. Valuation at any of the six rounds. Any confirmed frontier-lab customer — the Framer / Webflow / Lovable / Replit / HeyGen list rests on one launch newsletter and is uncorroborated. The nature of the Figma relationship — a "Case study: Figma" link exists on the launch post but the case study itself was unreachable. The names of the Head of Contra Labs, Head of Research, Data Lead and VP of Product — the roles are confirmed to exist by interview loops; the people are not named anywhere. Current headcount — three sources give three bands: 1–10 (Crunchbase), 11–50 (Clay), 51–200 (TheOrg). The size of the actual vetted Contra Labs Network, as distinct from 1.5M/1.7M marketplace registrations. Expert contract terms — no NDA or exclusivity agreement is public; absence of a clause in recruiting copy is not proof of absence in the project agreement. Any first-hand worker account — nothing on Reddit, Blind, Glassdoor or designer forums. And, notably for a company screen-recording professional sessions that contain client material: any legal pages at all — contra.com/legal and its variants 404 to automated fetch, and contralabs.com has no terms, privacy policy or DPA linked in its footer.
What to steal, and what not to
Start from captive supply or do not start. Acquiring 1.5M creatives cost the parent six years and $45M; it cost Contra Labs zero. If you do not own a vertical community, partner into one before building the data business.
Steal the phase decomposition. Converting a leaderboard into a diagnostic is the best idea in this company, and it is domain-portable.
Steal the trajectory capture — and lead with it. Narrated screen recordings of real client work, with execution_paths and preferred_execution, are the part a lab cannot synthesise or cheaply build in-house. Contra buried it under preference pairs.
Steal the creator-friendly terms. Licence not assignment, byline, no platform fee, seven-day pay, right to decline — cheaper than a 30% rate premium and far more persuasive to a wary profession.
Steal the funnel where the work test is the work, and the co-development partnership for the capability you lack: Design Crit exists because Contra rented Lica's modelling credibility instead of hiring an ML team.
Do not lead with a vanity network number. 1.5M against a 31-person study is a credibility liability with exactly the buyer you are selling to. Lead with rater qualifications and inter-rater reliability.
Do not give away eight datasets. Give away one — the one that proves the methodology — and gate the rest.
Do not build the eval as the business. Evals get you the meeting. Nobody has built a large business invoicing for them.
Do not confuse application-layer companies with labs. Size the real data budget of the real buyers in your niche rather than assuming the frontier-lab budget applies because the deck says "frontier labs".
Do not staff it as a pure services org and expect a product outcome, and do not inherit a loaded cap table: Contra Labs must clear roughly $45M of prior preference before common sees anything, which is an argument for a new entity even when the supply comes from an existing community.
The specialist wedge
The bet is that one domain buys cheaper experts and faster belief, and that both advantages expire the moment you have a reference customer. What would have to be true, what the evidence supports, and the trade-off that decides which niche.
Accounting, audit and tax
The cheapest credentialed pool in the set, the only one with a queryable national licence register, and nobody selling into it — against the second-thinnest evidence that a frontier lab wants it.
Eleven labour markets, not one
Freelancers bill 22.4 hours a week against a 50–60% utilisation benchmark, so hours 23–40 earn zero — which means the correct price for a bench hour is at market, not above it. Recruit the majority-self-employed segments, not the highest-paid ones, and expect to clear at $60–85/hr with no premium for creative expertise.
Rebuild the Human Creativity Benchmark
Take the flagship benchmark of the company being modelled, rebuild it properly, and make the contribution the reliability methodology rather than the leaderboard. Nobody has published how many designers per item an aesthetic ranking needs before it stabilises, by criterion — Google says more than ten and Contra's published rubric says three. That sweep is a re-analysis, not a new field, and it is the single most valuable unclaimed artefact in the space.
Taste Labs, in full
The dangerous competitor is not Contra. It is the company whose founder already sold to foundation labs, whose fourth and fifth hires train models, whose raters nominate each other — and which has published a research agenda promising to build the one asset Contra actually owns.
The taste read
The lane is chosen, so the question is what is true about it. Mercor already sells the premium version at $150–250/hr, the arena layer holds 6,047,075 image votes collected free, the data costs $0.91 a comparison, and a 10–25x pricing contradiction sits unresolved at the centre. Three things are genuinely in the client's favour and one of them is a property right no US competitor can hold.
What nobody has built
Eight absences, ranked, each searched specifically and each with what it would take and what it would earn. Zero design reward-model weights exist from any of the six organisations examined; no design benchmark anywhere is held out; nobody has run designer-panel against crowd-vote on the same items though both companies whose business models depend on the answer have the data; and the trajectory position is now contested.
What publishing actually bought them
Thirty-nine studies, two papers, eight datasets and no GitHub organisation bought Contra Labs 5,133 all-time downloads and zero citations. Patronus out-published them on their own moat in fourteen days; AfterQuery's 148-row FinanceQA has 17 citations and a repo. Volume bought nothing, and the pattern that did work is small, sharp, cited and shipped with code.
Design Arena, in full
The $60M ARR figure is self-reported, has no consumer revenue behind it, and is contradicted by the round it was announced with — a company at $60M does not raise a $7.9M seed from Index. What is real is 5.3M people voting for free, which is Pick-a-Pic commercialised.
Law
The largest credentialed pool anywhere in this atlas, a real observed clearing price of $140–160/hr, no vertical specialist — and a privilege problem with the cleanest workaround in the report.
Taste Labs
$18.5M from CRV and Amplify to sell judgment. Design is the beachhead, not the business — which is the difference between this and everyone else in the lane.
The creative tool layer
The 'probable first customer' hypothesis, tested against 1,562 job postings and largely failed: one dedicated AI-quality-evaluation role, zero named vendors, and $31.6B of vibe-coding valuation employing nobody to evaluate design quality. Canva is the exception, says so in writing, and has already built its own supply.
The eleven problems a rival published
Taste Labs published eleven open problems on 16 August 2026, the same day it opened a fellowship. It is a hiring instrument and it is also a public statement that a funded rival has not solved any of them — two of which are Contra's live products. Answering a named competitor's published open problem is the cheapest citation available in this market.
What to build first
Taste has no oracle, so inter-rater agreement is the manufactured one — which makes the panel, not the file, the product. Build a calibrated, named, consented panel and sell the sealed evaluation content that falls out of it; run it first in product and UI design on de-novo briefs; and remember that the corpus is a $1M line item a lab insources in one meeting.
Design Arena
5.3M users, $60M ARR, under a year old — the consumer-arena route into the same data. Nobody has said whether that number is gross or net.
The arena layer
Six million image votes collected free, $100M annualised on 28 people, and 80% of the preference data retained as a strategic asset. The specialist loses this fight on volume permanently — and Yupp is the corpse that proves both that generic preference is unsellable and that the demand side is thinner than the supply story implies.
Will they say yes
Stated attitude is 81%, revealed behaviour is 10%. The Sora revolt is the strongest evidence that creatives object to being unpaid, uncredited and approval-gated rather than to evaluating AI output — every demand was a term sheet, not a principle. But 51% object to who profits, not to being underpaid, and money does not fix that.
AfterQuery and UI-Bench
The most rigorous public human-judged design benchmark was built by an expert-data company a year before Contra Labs existed, from 194 hand-picked experts who appear to have been unpaid. Publishing a benchmark for credibility is not a novel move here — it is the category's standard customer-acquisition tactic.
Senior code review
Writing code is saturated and over-served; judging code has near-zero benchmark coverage. The only public, verifiable credential in this atlas — GitHub review history — sits inside the most crowded lane.
Telling a good designer from a confident one
There is a published pass mark: human experts hit 68.9% on the Visual Aesthetic Benchmark where the best frontier model manages 26.5%, and direct ranking yields substantially higher inter-annotator agreement than score-derived ranking. Collect rankings, screen at roughly 69% agreement with expert consensus, and treat peer nomination as a seeding mechanism rather than a scaling one.
Clinical medicine
The best-documented lab demand in the set and the widest open benchmark among the professional domains — attached to the highest cost basis anywhere, where a gastroenterologist's opportunity cost exceeds Mercor's entire ceiling.
The generalists in creative
Mercor already sells the exact product — narrated senior design reasoning — at $150–250/hr, from Pentagram and Wolff Olins pedigree, to an unnamed frontier client. Contra's 'up to $100/hr' is the middle of the band, not the top of it.
Halluminate
Occupies the finance niche by name, with $160K of disclosed funding that is almost certainly stale. Whether investment banking is taken or barely touched turns on a number nobody has.
Centaur Labs
The clinical incumbent: $31M raised to crowdsource medical annotation through a diagnosis game. It sells labelled artefacts, which leaves physician reasoning unsold.
Design and UI/UX
The thesis that started this whole enquiry, proven by three funded companies that between them already hold the network, the rubric science and the preference collection.
Video, motion and VFX
The only niche in this atlas with a frontier-lab rate card naming the exact tools, a public benchmark at 34.65%, and zero specialist competitors — attached to the smallest expert pool here, which is the constraint and the moat at once.
Game development and 3D art
A rich pool with excellent free channels, sitting behind a platform holder who has already banned the thing you would need to do, in an occupation BLS projects at 0% growth.