Miju Labs

The specialist wedge

The bet is that one domain buys cheaper experts and faster belief, and that both advantages expire the moment you have a reference customer. What would have to be true, what the evidence supports, and the trade-off that decides which niche.

medium confidence10 minupdated 2026-08-30strategy · niches · specialisation · benchmarks · expert data
Two of these domains have since been taken all the way down

This page argues the wedge in the abstract across sixteen niches. Two have since been researched to the bottom and written up as full dossiers, and both verdicts complicate what is below. Security scores 20 and still says build, on a narrower shape than the screen described, with the six firms named and a direct competitor inside the position — The security read. Health scores 24, the highest in the atlas, and the screen's confident cost objection turned out to apply to one specialty out of thirteen — The health read. Where a claim here conflicts with either, the dossier is the current reading.

The bet has a precise shape, and it is not "niches are good."

A vertical specialist claims exactly two advantages over the horizontals of Expert data for frontier labsMercor, Surge AI. It reaches better experts more cheaply, because a GREM-certified malware analyst or a Local 700 colourist answers a company that speaks their language and ignores a generic annotation ad. And it is believed faster, because a buyer judging a supplier in an unfamiliar domain cannot assess the data and can assess the benchmark you published about it.

Both advantages are front-loaded, which is the part the thesis usually skips. Cheap reach matters while you are assembling the first few hundred experts and stops mattering once you can pay market rate out of revenue. Fast belief matters before you have a reference customer and is worth roughly nothing after — the second lab buys on the reference, not the benchmark. The specialist advantage is a financing mechanism for the cold start, not a moat. See Which side you build first.

So the bet pays only if four things hold at once. The domain's best experts must be unreachable by an ad budget — a register, a leaderboard or a guild a generalist will not bother to work. There must be an unclaimed public artefact, a benchmark or one embarrassing number, that the buyer will actually read. Somebody with money must already be buying here, or be capable of being made to. And you must convert belief into a signed contract before a horizontal with existing lab relationships notices the category exists.

Fail the first and you are a worse Mercor. Fail the second and you have no cheap way in. Fail the third and you have built a research lab. Fail the fourth and you have done a funded competitor's market research for them.

The example is not what it looks like

Most people arrive here because of Contra Labs, which is not the thing it is being used to prove.

It is not a startup. It is a business line of Contra.Work Inc., a Delaware corporation incorporated in 2018 that runs a commission-free freelance marketplace (SEC Form D). The Labs unit launched 31 March 2026, five months after a $740,000 cheque from Zentavo VC, against $44.5M raised at a Series A and B back in 2021 (Clay). Five open roles, zero of them ML — the one modelling artefact was co-built with an outside partner (Ashby; Design Crit). The only corroborated customer list — "Framer, Webflow, Lovable, Replit, HeyGen" — rests on one launch newsletter and is entirely application-layer companies, not model labs (Creator's Toolbox) [UNVERIFIED]. No frontier lab is named as a customer anywhere. The full account is on Contra Labs.

None of that kills the thesis. It relocates it. Contra Labs demonstrates that an existing vertical community can be repointed at AI data quickly and credibly — 39 studies and an arXiv paper in five months bought a new unit more standing than a year of outbound would have. It does not demonstrate that a frontier lab will pay a design specialist, because no lab has been shown to. Anyone thinking "this works, do it in another domain" is reasoning from a company that has not yet proven the thing they want proven, and whose supply advantage cost a different business six years and $45M to build.

What the evidence actually supports

For. Three companies show the wedge working, and each shows a different mechanism.

Vals AI raised $40M at a $400M valuation led by a16z in August 2026, revenue reported up 8x on all of 2025, benchmarking law, banking, engineering and medicine, with results appearing in model cards from OpenAI, Anthropic, Google, Meta and xAI (Pulse 2). It entered professions the labs were not yet buying for and the spending followed. Gray Swan AI is cited in eleven frontier model system cards, with 20+ customers and an Arena of 15,000+ researchers (National Law Review). Irregular runs cyber evaluations for four labs with about 35 people at $80M raised on a $450M valuation, its CyScenarioBench named in the Claude Opus 5 system card at 33.7% completion (BERI; Anthropic).

Thirty-five people, four of the world's most sophisticated buyers, $450M. The wedge at full extension.

Against. Two negatives, both load-bearing.

The two negative findings

No frontier lab job posting anywhere hires lawyers, accountants or salespeople to produce data. Labs hire those professions for in-house functions only. The domains carrying named data-producing reqs with pay bands are cyber — OpenAI's Red Team Specialist at $198K–$320K, responsibilities explicitly including "constructing datasets" — health, and investment banking (OpenAI; Anthropic).

No system card in any of the sixteen domains explicitly cites purchased expert data. The strongest lab-side evidence in this atlas is job postings, rate cards and benchmark methodology sections — not a disclosed contract. Every "a lab is buying" claim below cyber and clinical is an inference from an intermediary's price list.

Hold those together. The wedge works in security, where the buyer publishes its vendors by name. Everywhere else the demand is inferred from what Mercor advertises it will pay — $140–160/hr for senior lawyers, $95–170/hr for civil engineers, $65–80/hr for mechanical engineers — which proves someone is buying, at an unknown price, from somebody who is not you.

The trade-off that decides everything

One finding organises all sixteen niche pages, and it is uncomfortable: the demand signal and the benchmark gap are inversely correlated.

NichebudgetproofRead
Offensive security52Named in system cards; benchmark layer vendor-captured
Senior code review53SWE-bench Verified saturated at 87.6–96%; twenty vendors
Investment banking and financial modelling53Lab keeps in-house SME headcount; Vals owns the benchmark
Clinical medicine54HealthBench and MedHELM open; physicians cost $132.66/hr
Law34No drafting benchmark exists; no lab hires lawyers for data
Accounting, audit and tax24No public audit benchmark at all; no lab hires accountants
Defensive security35Meta and CrowdStrike publish 23–34% and no business exists
Video, motion and VFX45The one place both hold at once

Where labs demonstrably spend, vendors already own the measurement layer: Gray Swan's Arena, Irregular's SOLVE and CyScenarioBench, Scale's SWE-bench Pro. Where the benchmark layer is empty, nobody has proved a buyer. You are choosing between fighting for money that exists and manufacturing money that does not.

Vals AI is the existence proof that manufacturing works — and simultaneously the reason the manufacturing lane is closing. Every quarter it operates, one more profession stops being unclaimed. Manufacture a budget in law or accounting and you are racing a funded company already standing on that step in four adjacent domains, which could extend downward into supply with one hire.

The six scores, and how they lie

Every niche page carries six scores, one to five, higher always better for the operator — the How to evaluate one of these's six re-cut for a domain rather than a market. The question is no longer "is this a good business" but "would this domain be a good first customer set."

ScoreThe questionA 5Where it lies
proofCan you become visibly the best source here inside two quarters?An unclaimed benchmark, a number to beat, an audience that reads itAn empty benchmark field with no audience scores 2 — Skilled trades and field service and Sales and GTM have zero competition and nobody to be best in front of
budgetHas anyone with money actually bought in this domain?A lab req with a pay band, or a name in a system cardIt measures the buyer's cheque, not your access to it. Most of it arrives via a middleman's rate card
defenseDoes the position survive a competitor and three years?A decade-long credential plus data that decays on a scheduleIt scores the domain, not your position in it. Life-science wet lab scores 4 and you still lose it to $720M of competing capital
costWhat does one unit of data cost — labour floor plus capital?No rig, no licence, no building, and a cheap credentialed hourA low wage floor can be a trap: the seller who will annotate at $95/hr is, by revealed preference, not a top performer
reachCan you find and verify the experts?A queryable national register — NASBA's 650,667 licensed CPAs, state bar admission, NCEESThe least informative score here. Sales and GTM, Skilled trades and field service and Game development and 3D art all score 5 and are all avoid
roomIs the position actually empty?No specialist in either vendor directoryEmpty is usually a finding. Seven empty categories means the buyers are absent, not that seven founders missed one idea

Two disqualifiers override any total.

What would kill it

An entrenched specialist caps room at 2 however good the rest looks. Design and UI/UX scores room 1 against Contra, Taste and Design Arena; Life-science wet lab scores 1 because Mercor bought Sepal AI in February 2026. You do not out-execute a funded incumbent in a domain whose entire buyer pool is a dozen accounts — that is One customer is a binary event working against you.

Data that needs a rig, a licence or a building before the first invoice. Chip design and EDA needs $80–150K per engineer per year in EDA licences; Life-science wet lab's marginal experiment costs hundreds to thousands of dollars; Skilled trades and field service needs a shop, instrumentation and insurance against a comparable market pricing this data at $1/hr (TechCrunch). Capital cost with a labour-cost buyer is not a business.

The totals are in the rank table. Read them as an invitation to argue — the six are correlated, and summing correlated judgements produces a number that looks like arithmetic.

What the shortlist is

Three survive, each for a different reason.

Video, motion and VFX is the only niche where the trade-off above does not bind. xAI's AI Tutor – Video Specialist pays $40–75/hr and names Premiere, Resolve, After Effects and Nuke; GDPval includes film and video editors; VEBench puts Gemini-2.5-Pro at 34.65% on editing-technique recognition, which its authors call "far below the level required for practical editing" (arXiv 2605.03276); and there is no specialist in either vendor directory. A rate card, an embarrassing number and an empty field is a combination that appears once.

Mechanical CAD and manufacturing is the best risk-adjusted second. The buyer already pays $65–80/hr through Mercor, the vertical is empty, GrabCAD's 16 million members is the largest verified-by-publishing channel here, and the dominant tool vendor is provably not hoarding the way Cadence is. The business is repricing mechanical engineering from generic STEM reasoning to a vertical.

Defensive security is the best manufacture bet, and the only one where the empty category is corroborated by a lab publishing the gap rather than by silence. Meta and CrowdStrike built CyberSOCEval and stated that models are "far from saturating" it, at 23–34% on malware analysis (arXiv). There is no SWE-bench for the SOC. The constraint is that you never touch customer telemetry: the product is a certified analyst narrating their triage of a public malware sample, never a log file.

Proven but taken. Offensive security, Senior code review, Investment banking and financial modelling and Life-science wet lab carry budget 5 or its intermediated equivalent and room 2 or less. Sell to the incumbent rather than compete with it. Design and UI/UX belongs here too: the inspiration for the model is the one market you should probably not enter.

Clear no. Sales and GTM and Skilled trades and field service on absent demand — when xAI enumerated the four domains it was surging specialists into, it named STEM, finance, medicine and safety (TechCrunch). Chip design and EDA and Game development and 3D art on structure: the tool vendor owns the data layer in one, a platform holder banned the necessary activity in the other.

Accounting, audit and tax, Law, Clinical medicine and Architecture, engineering and construction sit just below the line, each on one problem — no proven buyer, no proven buyer, a $132.66/hr wage floor, one observed purchase. Solve that problem and any of them displaces something above it.

What would settle this

None of the above is worth another week of desk research. Five things would move these scores further in a fortnight than the atlas did.

Pull one register and measure the conversion. NASBA's Accountancy Licensee Database holds 650,667 actively licensed CPAs across 53 jurisdictions (NASBA) — the only queryable national register in the set. Sample 300, spend a small fixed sum, count qualified applicants in a week. That number is the reach score, and if a register with a URL does not convert, no channel-based niche will.

Go to one conference and say the offer out loud. NAB or a Local 700 event for Video, motion and VFX; BSides or BlueTeamCon for Defensive security. The question is not whether editors exist. It is whether a workforce whose expectations were set by Adobe's $4.88 Firefly bonus takes $75/hr to critique AI output, or spits. No subreddit answers that.

Have three conversations, and only these three. One person lab-side who has actually signed a data vendor contract, to establish whether a vertical line item exists at all or whether everything routes through a horizontal. One seller who has already sold a vertical dataset — Vals, Halluminate, the Rivet team behind TaxBench — asked one question: how did the published benchmark convert into a purchase order. And one application-layer buyer: Fieldguide, Basis, Rogo, Leo AI. That third is the Contra Labs lesson operationalised — if the real buyer is a Series C app company, the addressable spend is one to two orders of magnitude smaller and the plan changes shape.

Reproduce one benchmark. Take VEBench's 34.65% and re-run it on fresh long-form footage with a current model. If it replicates, you have your one sales statistic — the equivalent of Contra's line that off-the-shelf VLM judges reach 54% agreement with designers against a 74.1% human ceiling (Design Crit). If a current model scores 70%, the proof score collapses and you have saved yourself a company for a week's work.

Get one price. Ask anyone selling a vertical dataset for a quote.

The hole under all of it

No pricing data exists for any vertical data contract in any of these sixteen domains. The only public anchors are ~$20,000 for a website replica and per-hour expert rates. Nobody publishes what a CAD eval set or a VFX preference dataset sells for — and where revenue figures do exist they are gross payment volume, not net (GMV is not revenue). The number that would decide this business is the one nobody has. See What we could not establish.