Miju Labs

All specialists

Contra Labs

Not a startup — a business line of a six-year-old freelance marketplace, launched five months after a $740K cheque. The template everyone wants to copy, and the parts of it that do not survive inspection.

medium confidence11 minupdated 2026-08-30design · creative data · preference data · trajectories · evals
Vertical
Design
Founded
Contra.Work Inc. incorporated in Delaware, 2018; Contra Labs launched 31 March 2026
Headquarters
San Francisco on the SEC filing; the Labs unit is being built in Williamsburg, Brooklyn
Raised
~$45.2M at the parent, across all rounds. No Contra Labs–specific round exists
Last valuation
Not disclosed, at any round
Revenue
Not disclosed
Status
Active — five months old, five open roles, no named customer
Who runs it · 4 people in the index

Latest (Sep 2026): Contra launched Contra Labs on 31 Mar 2026 as 'the first frontier data and evaluation lab for creative AI'. It draws on a claimed 1.5–1.7M-member creative network and offers a Creative Arena, the Human Creativity Benchmark, preference data and screen-recorded trajectories. Research followed in June: Design Crit (17 Jun; off-the-shelf VLM judges below 55% agreement with designers, a small trained head at 61.1%), the TASTE dataset co-built with Lica World, and the HCB paper on arXiv (29 Jun). Its modelling partner Lica was acquired by Gamma on 25 Aug 2026. In Aug 2026 it was hiring Research Scientists (Human Data & Evaluation) and project leads at $150–200K in NYC and SF. No frontier-lab customer has been named publicly.

What Contra Labs is saying
Contra
67,517 followers
What happens when you give ChatGPT 'the worst café branding brief imaginable'? Chhavi handed it rough sketches and a loose idea, and in under 30 seconds, ChatGPT Images 2.5 turned it into something that actually felt cohesive 🔥
114 comments2 reposts
Contra
67,517 followers
Love seeing creators on Contra put new tools to work! Last week, OpenAI dropped ChatGPT Images 2.5 with Sketch, a feature that turns rough drawings into complete images. Ana ran it on a 30-second doodle and the output was not what she expected 👀
181 comments2 reposts

Contra Labs is not a startup. It is a business line of Contra.Work Inc., a Delaware corporation incorporated in 2018 that runs a commission-free freelance marketplace (SEC Form D). The homepage says so in the footer — "Powered by Contra" — the LinkedIn link points at linkedin.com/company/contrahq, and the jobs board is jobs.ashbyhq.com/contra with Contra Labs appearing as a department inside it (contralabs.com; Ashby posting API). EDGAR full-text search returns zero hits for "Contra Labs" across all form types (EDGAR FTS). There is no separate legal entity.

That correction reframes what is actually being copied. The launch was 31 March / 1 April 2026 (Contra blog; Creator's Toolbox, 1 Apr 2026) — five months after a $740,000 venture cheque from Zentavo VC in October 2025 (Clay funding dossier). Before that: a $14.5M Series A led by Unusual Ventures in February 2021 and a $30M Series B led by NEA in November 2021, with Cowboy Ventures participating (TechCrunch, Feb 2021; TechCrunch, Nov 2021). Total raised is about $45.2M (Clay).

A company that raised $30M in 2021 and takes $740K four years later is not on an up-and-to-the-right path. The funding shape says the marketplace stalled and this is the second act — an incumbent repointing existing supply, not a new company assembling one. Every "how did they get 1.5M designers so fast" question has the same answer: they did not. It took six years and $45M, paid for by a different business.

A research trap — do not cite this page

The Crunchbase profile at crunchbase.com/organization/contra-labs is a different, defunct company: a Brazilian app developer founded in 2007 in Salvador, Bahia by Victor Cardozo and Gustavo Carvalho, listed as permanently closed, whose contact address was contato@contralabs.com (Crunchbase). Contra evidently acquired the domain from them. The correct parent profile is Crunchbase — Contra.Work Inc. Anyone building a competitive map from Crunchbase alone will import a dead Brazilian company's founding date into their model of the Design and UI/UX vertical.

What they actually sell

Three commercial lines and one marketing line, and the ordering on the homepage is backwards.

Preference pairs and rubric scores are the headline product: "Seed datasets (image, video, UI, motion). Corrective SFT data and RLHF preference pairs" (contralabs.com). The flagship study alone produced 5,940 pairwise judgements and 3,675 written rationales, plus 1–5 Likert scores on Prompt Adherence, Usability and Visual Appeal (arXiv 2606.30561). The public rubric is five-axis — visual quality, prompt adherence, originality, utility, and motion realism for video — with three or more professional evaluators per output (contralabs.com/human-creativity-benchmark).

Design Crit is the reward-model line. "Criteria-Resolved Image Taste" is a designer-annotated preference set scoring nine dimensions rather than one verdict, on which they trained "a small pairwise-difference head on top of a frozen vision-language encoder, with no fine-tuning of the backbone" — their words, a "deliberately modest model" (Design Crit). It is explicitly "a Lica × Contra collaboration": Lica World is a separate creative-AI research company positioning itself as "built for research labs developing the next generation of design AI" (lica.world/blog). The modelling capability was partnered in, not built.

Screen-recorded trajectories are the differentiated asset, and they are buried third. The Premiere Pro dataset card gives the shape: 234 steps across 4 trajectories, 111–245 minutes per session, recorded on macOS while professional editors built vertical social reels from real client briefs. Per-step schema: trajectory_uuid, session_uuid, image, thought, action_type, tool_call as structured JSON, execution_paths (MCP tool / keyboard shortcut / menu path) and preferred_execution (HF card). The card's claim is the whole pitch: "The screenshots, the recorded action, and the pointer coordinates are the editor's real execution" — and the thought field comes from the editor's spoken narration, explicitly distinguished from "model-synthesized rationales".

That last detail is the one to take. You cannot synthesise a professional's mouse path through Photoshop, and you cannot obtain their reasoning without them talking while they work. Preference pairs are a commodity — any lab with thirty contractors and two weeks makes them. The narrated trajectory requires a standing relationship with professionals doing real client work, which is exactly what Which side you build first says money cannot buy quickly. Contra owns the moat and leads with the commodity.

Read the dataset names carefully

gemini-creative-campaign-trajectories and firefly-creative-campaign-trajectories are named for the tool the designer used, not the customer (HF org). They are not Google or Adobe deals.

The phase decomposition, which is the smartest thing here

The Human Creativity Benchmark splits creative work into three phases — Ideation → Mockup → Refinement — and shows that model rankings invert between them. Claude leads Ideation on landing pages, Gemini dominates Mockup at a 68.9% win rate, Claude reclaims Refinement (HCB). Veo 3.1's 61% ideation win rate falls to 39% at refinement (contralabs.com/research).

The methodological claim underneath is that evaluator disagreement in creative domains is signal, not noise — it separates convergence on shared professional standards (typography, layout, hierarchy) from divergence in legitimate taste (arXiv). Agreement is highest on Prompt Adherence and lowest on Visual Appeal, as you would expect if the claim holds.

This matters commercially, not intellectually. A leaderboard tells a lab "you are third", which is useless and unwelcome. A phase decomposition tells them "your model collapses at refinement, here is why", which is a purchase order. It is portable: every domain has an analogous split — issue-spotting → drafting → redlining in Law, differential → workup → management in Clinical medicine. Finding yours is the transferable move; see The specialist wedge.

The buyer gap

The job posts describe the customer as "frontier AI labs and product companies" and "frontier AI research teams" (Ashby; Ashby). No frontier lab is named as a customer anywhere — not OpenAI, not Anthropic, not Google, not Meta, not xAI. There is no logo wall, no case study page, no testimonial.

The only corroborated partner list comes from a launch newsletter, verbatim: "Contra Labs is already partnering with a wide range of AI-powered creative tools (including Framer, Webflow, Lovable, Replit, HeyGen, and others)" (Creator's Toolbox). [UNVERIFIED] — it could not be corroborated from Contra's own site, and "partnering with" spans everything from a signed contract to a shared Slack channel.

Every name on that list is an application-layer AI company, not a model lab — Series B/C startups with product budgets in the tens of thousands, not labs with nine-figure data budgets. Mercor's revenue comes from OpenAI, Google DeepMind and Meta (Sacra). If Contra Labs' real buyer base is design-adjacent app companies, the addressable spend is one to two orders of magnitude below what the "Mercor for design" framing implies. That is the most important risk here, and the thing to stress-test first in your own niche — see How much money is actually in the buyer pool.

The 39 published studies evaluate Claude, Gemini, GPT, Qwen, Veo, Seedream, FLUX, Kimi, Grok, Nano Banana and Meta Muse — as subjects, not as clients (contralabs.com/research). "Where four AI models break when they build a landing page" is a lead-generation tactic aimed at those labs, not evidence of a relationship with them.

The supply terms — the IP is the weapon, not the rate

TermWhat Contra Labs offersSource
RateUp to $100/hr, depending on role and seniority; some briefs flat-feecontralabs.com/jobs
Brief lengthTypically 5–10 hours, occasional longer engagementsHelp Center
PaymentWithin 7 business days of task completionHelp Center
Platform feeNone taken from the expert's ratecontralabs.com/jobs/photographers
IPExpert licenses work to Contra Labs — not an assignmentcontralabs.com/jobs/photographers
CreditPublished work carries an author bylinecontralabs.com/jobs/photographers
RefusalExperts "can decline projects that don't feel appropriate"contralabs.com/jobs/photographers
ExclusivityNone found in any recruiting page, help article or job post [WEAK on the negative]

The rate is not the recruiting weapon. Up to $100/hr sits above Mercor's $85+/hr average and well above micro1's $30–65/hr creative band (Sacra; aitraining.jobs) — but it is below what a good brand designer bills a real client. On price alone, Contra Labs is fill-in work.

The IP terms are the weapon. Licence-not-assignment, byline credit, no fee skimmed off the rate, the right to decline, seven-day pay instead of thirty to sixty. In a profession ideologically wary of training the thing that replaces it — Envato's survey of 1,780 creatives found graphic designers and illustrators have the lowest daily AI adoption at 40% and the highest frustration (Envato) — those terms are cheaper than a 30% rate premium and considerably more effective. The intake funnel matches: portfolio review, a video interview compressed in one role to a 30-second intro, and an assessment that is "evaluating AI-generated brand assets and providing structured feedback", about 10 minutes in total (Help Center). The work test is the work.

None of it buys lock-in. No exclusivity clause appears anywhere in the public funnel; experts keep a licence, can decline, and can take Mercor's rate tomorrow — the standard Getting cut out exposure of a marketplace that does not own the relationship.

The scale claim collapses on contact with the methodology

The homepage says "1.5M+ verified creative experts" in the hero and "1.7M+ creative experts" in the stats block, on the same page (contralabs.com). The benchmark page repeats the contradiction (contralabs.com/human-creativity-benchmark); the Creative Human Data page says 1.7M+ alongside "400+ skills and tools" and "$250M+ verified expert earnings" (contralabs.com/creative-human-data); the arXiv paper says 1.5M+ (arXiv).

Set either number against the studies. The flagship benchmark used 28–31 evaluators from 13 countries across 80 sessions, 93–95 prompts and 380 outputs from 13 models (arXiv; HF HCB card). Design Crit used 10 professional designers in two cohorts of five (Design Crit).

The 1.5M counts registered profiles on the parent marketplace, not the vetted Labs roster — nothing claims 1.5M people passed Labs vetting, and the vetted number is never disclosed. A technically literate buyer spots this in the first meeting. Worse, it advertises how cheap the studies are to replicate.

The org chart tells you the ceiling

RoleLocationBandPosted
Strategic Project LeadNYC, onsite$150K–$200K + equity2026-06-08
Strategic Project LeadSan Francisco, remote$150K–$200K + equity2026-07-01
Special Project LeadNYC, onsite$150K–$200K + equity2026-07-30
Research Scientist, Human Data & EvaluationNYC, onsite$150K–$200K + equity2026-08-17
Research Scientist, Human Data & EvaluationSan Francisco, remote$150K–$200K + equity2026-08-17

Source: Ashby posting API.

Five roles, all delivery or research-operations. Zero ML engineers, zero infrastructure engineers, zero applied scientists who train models. The band is flat at $150–200K including the PhD-preferred Research Scientist — a quarter to a half of what a frontier lab pays research staff. One posting requires "data labeling or annotation experience": they are hiring from Scale, Surge and Appen, not from design studios. Another owns "case studies and benchmark publications", which makes the research programme a demand-generation programme by its own job spec.

The interview loops name an org chart nobody is named in: Research Scientist runs recruiter → Head of ResearchData LeadHead of Contra Labs → paid case study; Strategic Project Lead runs recruiter → CEO → technical → VP of Product → paid case study (Ashby; Ashby). The CEO personally interviews project leads — at five months old, Ben Huffman is running this himself.

Two reads. First, NYC Williamsburg, five days onsite, carved out of a parent that brands itself remote-first with no-meeting Tuesdays and Wednesdays (contra.com/careers) — the exemption a founder grants a unit that is the company's actual bet. Second, that org chart can only produce a consultancy. Delivery leads plus research-ops plus a rented model from Lica is an agency P&L with a data deliverable. If you want a reward model as a product, one of your first five hires has to train models.

Research as demand generation

39 studies in five months — roughly two a week — plus an arXiv paper and eight CC-BY-4.0 datasets on Hugging Face (contralabs.com/research; HF org).

DatasetRows / stepsDownloads (all time)Likes
HumanCreativityBenchmark8,012 rows5322
premiere-video-editing-trajectories234 steps (4 trajectories)1,41512
video-detail-annotation15 rows1,0073
photoshop-creative-design-trajectories294 steps8223
gemini-creative-campaign-trajectories266 steps4321
firefly-creative-campaign-trajectories137 steps3991
descript-video-editing-trajectories803 steps3800
creative-ad-design-dataset35 rows1462
A correction

An earlier version of this page reported far larger figures in the downloads and likes columns. Those were the datasets' row and step counts read into the wrong columns, and the argument built on them — that likes wildly exceeded downloads, so this was attention rather than adoption — does not survive the real numbers. The corrected figures come from the Hugging Face API and are reproduced in Every public artefact, by organisation.

The eight repositories total roughly 5,100 downloads all time, against 24 likes. The flagship benchmark — the artefact the company's whole research programme points at — has 532 downloads and two likes. The best performer is a set of four Premiere trajectories. For comparison, Patronus AI's FigmaTrace shipped 3,469 design trajectories in August and took 1,784 downloads in a fortnight. See What publishing actually bought them for what that gap means.

The one sales statistic

Off-the-shelf VLM judges reach roughly 54% agreement with professional designers. Chance is 50%. The human single-rater ceiling is 74.1%. Contra's trained head reaches 61.1% (Design Crit).

That is the commercial argument in one line: your LLM-as-judge is barely better than a coin flip on design quality, we have the humans who fix it, and a deliberately modest model proves the data moves the number. It works because it measures a deficiency in the buyer's own stack rather than asserting something about Contra. Find your equivalent statistic before writing a line of outbound.

Pricing

Undisclosed. No pricing page, no tiers, no "starting at", no rate card anywhere on contralabs.com or contra.com. The only commercial front door is a Calendly link — calendly.com/partnerships-contra/contra-labs-partnership-request — the primary CTA on every page (contralabs.com; contralabs.com/creative-human-data).

A Calendly, rather than a contact-sales form or self-serve checkout, is the tell of a business doing a single-digit number of bespoke deals. With five open delivery roles and a five-month history, buyer One customer is a binary event is near-total by construction. Any margin figure quoted for this company is an inference from the cost side; see GMV is not revenue before repeating one.

What could not be established

Pricing to any buyer — no rate card, no tier, no observed contract value. Revenue, ARR, bookings or gross margin — nothing public at either Contra Labs or the parent. Valuation at any of the six rounds. Any confirmed frontier-lab customer — the Framer / Webflow / Lovable / Replit / HeyGen list rests on one launch newsletter and is uncorroborated. The nature of the Figma relationship — a "Case study: Figma" link exists on the launch post but the case study itself was unreachable. The names of the Head of Contra Labs, Head of Research, Data Lead and VP of Product — the roles are confirmed to exist by interview loops; the people are not named anywhere. Current headcount — three sources give three bands: 1–10 (Crunchbase), 11–50 (Clay), 51–200 (TheOrg). The size of the actual vetted Contra Labs Network, as distinct from 1.5M/1.7M marketplace registrations. Expert contract terms — no NDA or exclusivity agreement is public; absence of a clause in recruiting copy is not proof of absence in the project agreement. Any first-hand worker account — nothing on Reddit, Blind, Glassdoor or designer forums. And, notably for a company screen-recording professional sessions that contain client material: any legal pages at allcontra.com/legal and its variants 404 to automated fetch, and contralabs.com has no terms, privacy policy or DPA linked in its footer.

What to steal, and what not to

Start from captive supply or do not start. Acquiring 1.5M creatives cost the parent six years and $45M; it cost Contra Labs zero. If you do not own a vertical community, partner into one before building the data business.

Steal the phase decomposition. Converting a leaderboard into a diagnostic is the best idea in this company, and it is domain-portable.

Steal the trajectory capture — and lead with it. Narrated screen recordings of real client work, with execution_paths and preferred_execution, are the part a lab cannot synthesise or cheaply build in-house. Contra buried it under preference pairs.

Steal the creator-friendly terms. Licence not assignment, byline, no platform fee, seven-day pay, right to decline — cheaper than a 30% rate premium and far more persuasive to a wary profession.

Steal the funnel where the work test is the work, and the co-development partnership for the capability you lack: Design Crit exists because Contra rented Lica's modelling credibility instead of hiring an ML team.

Do not lead with a vanity network number. 1.5M against a 31-person study is a credibility liability with exactly the buyer you are selling to. Lead with rater qualifications and inter-rater reliability.

Do not give away eight datasets. Give away one — the one that proves the methodology — and gate the rest.

Do not build the eval as the business. Evals get you the meeting. Nobody has built a large business invoicing for them.

Do not confuse application-layer companies with labs. Size the real data budget of the real buyers in your niche rather than assuming the frontier-lab budget applies because the deck says "frontier labs".

Do not staff it as a pure services org and expect a product outcome, and do not inherit a loaded cap table: Contra Labs must clear roughly $45M of prior preference before common sees anything, which is an argument for a new entity even when the supply comes from an existing community.