If the frontier labs are the prestige customer, the creative app layer is supposed to be the realistic one: dozens of companies, fresh capital, design quality visibly gating their product, and no in-house research organisation to build around. That is the hypothesis. It does not survive the evidence.
Across Figma (163 postings), Vercel (91), Lovable (78), Replit (71), Canva (274), Gamma (32), Webflow (27), Squarespace (24), OpusClip (11) and Descript (8) — 1,562 postings retrieved directly from Greenhouse, Ashby and SmartRecruiters APIs on 30 August 2026 — there is exactly one dedicated AI-quality-evaluation role, and zero postings naming any external evaluation vendor, design-review panel or expert-data supplier.
These companies employ designers in volume, at $105K–$376K. They do not appear to buy designer judgement.
$31.6B of valuation, zero evaluation headcount
| Company | Valuation | Design roles | Eval roles |
|---|---|---|---|
| Lovable | $13.3B — doubled in eight months on a $400M round | 10 | 0 in 78 postings |
| Replit | $9B — tripled in six months on a $400M Series D led by Georgian | 4 | 0 in 71 postings |
| Vercel (v0) | $9.3B, ~$200M ARR | 5 | 0 in 91 postings |
Between them, roughly $31.6B of valuation and nobody whose job is to evaluate design quality. UI-Bench ranks all three, and Replit sits at the bottom of it — a measured, public, unaddressed design-quality problem at a $9B company.
There is corroborating signal that the question is live: Business Insider covered "Can AI have taste?" as a hot topic at Replit's New York vibe-coding conference in June 2026. But interest at a conference is not a budget line, and no budget line is visible.
The most plausible reading is that these companies are growing so fast on distribution and model access that design quality is not yet the binding constraint. UI-Bench suggests it will become one — it exists precisely because outputs differ in design quality despite sharing underlying models. That is a timing argument, not a demand argument, and a client should price it as such.
Figma is pure build, and the eventual competitor
163 postings, 19 of them design, bands $105K–$376K, including Product Designer, AI Models and Product Designer, Design, Dev & AI Tools. Zero evaluation, human-data, annotation or rater roles. Its Designer Advocate roles at $153K–$317K are developer relations for designers, not evaluation (Greenhouse).
Figma is a company whose entire staff is a design-taste asset. It is the least likely buyer of external design judgement in the landscape, and the most likely eventual supplier of it.
Canva is the exception and says so in writing
Canva is the only app-layer company with a visible, funded, explicitly named evaluation function. Its Principal Research Scientist, Evaluations posting is the single most important document in the buyer research:
"Canva's generative models are judged by millions of people who will never read a benchmark. They just know whether the design looks right. Turning that judgement into something measurable is the hardest problem in our research stack, and it gates everything else. If we cannot measure design quality reliably, we cannot train against it, we cannot tell a real improvement from noise, and we cannot decide what ships." — SmartRecruiters
And, decisively for procurement:
"You will shape the principles teams use to trade off evaluation compute, human data spend..."
"Human data spend" is a named budget line at Canva. It is the only explicit confirmation of a human-data budget at any app-layer company in this research. The role is company-wide, "setting direction across our research groups in Australia, Europe, the US and China," covering design, image, video, audio and agentic workflows. Canva runs four other evaluation-adjacent roles: a second Principal Research Scientist for Evaluations in San Francisco, an Engineering Manager (ML) for the Evaluation Platform, and a Senior Research Scientist for Design Generation.
But Canva's supply answer is already built
The same board discloses how Canva sources the labour. Its AI Quality Evaluator (Dutch, 12-month contract, Amsterdam) is tasked with:
"Labelling, annotating, and evaluating photos, graphics, videos, stickers, and designs in Dutch — assessing linguistic quality, accuracy, and cultural appropriateness. Evaluating AI-generated content against Canva's quality bar for Dutch users... Collaborating with in-house Data Labelling teams across the Philippines and EU for ongoing quality alignment." — SmartRecruiters
So Canva has in-house labelling teams in the Philippines and the EU, per-language contract evaluators hired directly, and a Principal-level scientist allocating the budget between compute and human data. That is a sophisticated build, not a procurement gap.
Canva has already solved core design evaluation internally. What no in-house team scales to is language and culture: per-market evaluators assessing cultural appropriateness across dozens of locales, hired one contract at a time in Amsterdam and Seoul and Madrid. That is a staffing problem with a vendor-shaped answer, and it is the only opening on this page that a specialist can actually fill. It also happens to be a European-labour argument — see The one right a US competitor cannot hold and Eleven labour markets, not one.
One caution on timing: Canva's valuation was cut by roughly $10–11B in August 2026 from about $42B, against roughly $4B ARR, "putting IPO plans in doubt". A company defending a valuation cut is a harder sell for new discretionary spend and an easier sell for anything framed as replacing in-house headcount. Frame accordingly.
Buy versus build, company by company
| Company | Latest funding | Valuation | Design / eval hiring signal | Verdict |
|---|---|---|---|---|
| Canva | secondary/tender | ~$31B after Aug 2026 cut; ~$4B ARR | 24 design/eval roles; AI Quality Evaluator contract; 2× Principal RS Evaluations; in-house labelling in Philippines + EU | BUILD, aggressively — but the closest thing to a real buyer |
| Figma | public (IPO 2025) | public | 19 design roles, 0 eval roles in 163 | BUILD |
| Lovable | $400M (Aug 2026) | $13.3B | 10 design roles, 0 eval in 78 | BUILD |
| Replit | $400M Series D (Mar 2026) | $9B | 4 design roles, 0 eval in 71 | BUILD |
| Vercel | — | $9.3B, ~$200M ARR | 5 design roles, 0 eval in 91 | BUILD |
| Adobe | public | public | board not machine-readable (Workday); compensates Stock contributors by bonus, not evaluation contracts | BUILD + licenses content, not judgement |
| Webflow | $120M Series C (2022) | $4B (2022) | 27 postings, 0 design or eval roles | BUILD / not buying |
| Squarespace | Permira take-private, $7B (2024) | $7B | 1 design role in 24 | BUILD |
| Gamma | $68M (Nov 2025) | $2.1B, $100M ARR | 3 creative roles, 0 eval — and it just acquired Lica's design research lab | BUILD, now with a research arm |
| OpusClip | $20M, SoftBank Vision Fund 2 | $215M | 3 creative roles | BUILD |
| Descript | OpenAI-led round (2022) | $500M+ (2022) | 8 postings, 0 design/eval | BUILD / dormant |
| Framer | $100M Series D (Aug 2025) | $2B, ~$50M ARR | no public ATS | BUILD [WEAK] — inferred from absence |
| Wix, Bolt/StackBlitz, HeyGen, Captions/Mirage, Beautiful.ai | — | — | no public ATS | Could not establish |
| Tome | pivoted/wound down [WEAK] | — | no board | Not a buyer |
Boards returning zero creative or evaluation roles: Webflow (27), Descript (8), Ideogram (3), Pika (10), Hedra (8).
The Contra partner list could not be corroborated or refuted
The claim under test is that Contra Labs' only corroborated partner list is Framer, Webflow, Lovable, Replit and HeyGen, sourced from a launch newsletter.
contra.com/labsandcontra.com/aiboth return 404. There is no live partner page at the obvious URLs.- None of Framer, Webflow, Lovable, Replit or HeyGen mentions Contra, Contra Labs, or any external evaluation vendor across their 176 combined machine-readable postings.
- Webflow provides no signal either way; Framer and HeyGen have no public board at all.
So the list stays [UNVERIFIED]. What is corroborated is Contra Labs' existence, scale claim and pricing, via Fast Company on 24 August 2026: an "independent human data and creative evaluation lab launched by hiring platform Contra" connecting "more than 1.7 million creative professionals" with AI developers, where creatives can earn "from around $50 to $250 an hour working on RLHF and other types of post-training" (Fast Company).
Note the asymmetry in that rate: it is what workers are paid, which implies a materially higher price charged to buyers — typically 1.5–2.5× — but no margin or contract value could be established. GMV is not revenue applies.
What the silence actually means
Zero vendor mentions in 1,562 postings admits two readings, and honesty requires holding both.
Reading one: they are not buying. These companies employ designers, ship fast, and treat design quality as a staffing question. Every signal in the table above is a build signal.
Reading two: purchasing never touches a job description. This is common — procurement rarely surfaces in ATS text, and the same search across the frontier labs found no vendor names either, in a sector we know buys through vendors because OpenAI's own postings describe managing them. See The generative-media labs.
The methodologically defensible conclusion is narrower than either: there is no visible evidence of app-layer demand for external design evaluation, and the absence would look identical if the demand existed and was procured quietly. What tips it toward reading one is the corroborating structure — no evaluation headcount either, which purchasing would not hide. A company that buys external evaluation still needs someone internal to specify and consume it, and that person is not being hired anywhere except Canva.
Do not model the app layer as the first customer. Model it as: one qualified prospect (Canva, on localisation), one eventual competitor (Figma), and a cohort of very well-funded companies whose design-quality problem is real, measured, public and not yet anybody's budget line. The realistic first invoice is smaller and stranger than the hypothesis assumed — What to build first and Ninety days in taste work from that premise.
jobs.ashbyhq.com/runway is a business-planning startup, not RunwayML — its postings describe a product that "replaces traditional spreadsheets with a modern planning platform." RunwayML has no machine-readable board. Any competitor or buyer analysis citing "Runway" job-board data is almost certainly counting a different company's hiring, and the error is invisible once it reaches a spreadsheet.
Whether app-layer companies buy through unposted channels; Framer, Wix, HeyGen, Bolt and Captions have no public board at all; Adobe's is not machine-readable; and Contra Labs' actual contract values and partner list remain unestablished. Also unestablished: whether Canva's "human data spend" is six, seven or eight figures.