Miju Labs

A living index of the expert-data market · 15 markets · 16 niches · 3341 posts archived

The buyers have money
They do not have people

Fifteen markets where someone stands in the middle, assembles a supply nobody else can reach, and keeps the difference. The survey is done and the question has narrowed: one market, and which single domain to own inside it.

And it does not stand still: the companies, the people inside them and what they are saying, indexed as it happens.

00

A living index of the field

The research below was a snapshot. This is becoming the running record of the whole market: every company in it, every person worth knowing inside those companies, and what all of them are doing and saying — updated continuously, archived permanently, and cross-linked to the analysis.

Latest from the field
Vals AI
10,294 followers
AI models are advancing faster than many of the benchmarks meant to measure them. TechCrunch visited our office to look at how we’re building a more neutral, rigorous and trustworthy evaluation layer for AI. Our work goes beyond abstract tests of intelligence to examine whether models can perform real-world work — and what risks emerge when they are deployed — across myriad industries including law, finance, coding, cybersecurity, biosecurity, and mental health. Thank you Lucas Ropek for spending time with our team and telling our story.
1
Rayan K.
Evaluating LLMs with Vals AI
Independent evaluation is only credible when evaluators have meaningful access, real independence, transparency about the terms of their work, and the freedom to publish their findings. I signed the AI Evaluator Forum’s statement because these principles are closely aligned with how we think about evaluation at Vals AI. If embedded evaluation is going to strengthen trust in frontier AI, the conditions have to be right.
31 comments
01

Where this went

Fifteen markets became one — expert data for frontier labs — then sixteen candidate domains became three dossiers, and the lane is now chosen: creative work and taste. The earlier rounds are kept whole rather than tidied away, because the reasoning is the point.

The taste read · the build plan · the people map

All 16 niches · the 9 companies that already picked one · the Contra Labs dossier

02

Six questions decide everything

Every vertical page answers the same six questions in the same order, so the answers sit side by side. Four of them do most of the work.

01 · BUD

Budget depth

How much money sits behind the buying decision, and how urgently it must be spent.

02 · SUP

Supply difficulty

How hard the seller side is to assemble. Hard is good: it is the only thing a competitor cannot copy in a weekend.

03 · SPR

Spread

What fraction of the money passing through you, you actually keep.

04 · HLD

Holdability

Whether buyer and seller can cut you out once they have met, and whether one buyer can end you.

05 · AI

AI direction

Whether capable models grow this budget or delete it.

06 · SPD

Speed to first dollar

How long from a standing start to a real invoice.

The rubric, and the argument for why these six and not others, is in the framework. The name for the mechanism — and why it is not a marketplace, a staffing firm or an agency, though it is sold as all three — is in what the model is.

03

Where the spread actually survives

Ranked by the six scores. The sum is a blunt instrument and the ranking is a judgement, not a measurement — but the order is where the argument starts.

01
Adversarial evals and red-team crowds
A brand-new budget line with no salary to be benchmarked against, supply that cannot be recruited by job ad, and a middleman that monetises the crowd's output as software rather than reselling its hours.
26/30
build
02
Forward-deployed engineering
Deployment is the acknowledged bottleneck and the budget is real; the open question is whether any of these companies converts field work into a renewing licence before the 58x mark has to be justified.
22/30
watch
03
Robotics teleoperation and physical-world data
The only human-data vertical where physical capital blocks the laptop-marketplace playbook — which is why quality teleop still holds a 40–65% margin while egocentric video went to free in eighteen months.
21/30
build
04
Expert data for frontier labs
The budget is real, urgent and growing. The margin is a staffing margin, the buyers are two companies wearing five names, and every seat at the top table is taken.
20/30
crowded
05
Expert networks
A 70–80% take that survived forty years, three compliance scandals and every disintermediation attempt — because the product was never the expert, it was recruitment speed plus indemnity. Whether the incumbents sell to AI labs is the single most valuable unanswered question in this atlas.
20/30
open
06
Data-centre labour and site brokerage
The only capacity vertical where the unit is not fungible and not indexed, so a real spread can still live there — but the incumbents are ordinary staffing firms at 30–40% markups and nobody has built the AI-native version.
20/30
watch
07
Contingency recruiting marketplaces
The supply-side wedge is real and the AI tailwind is real; the TAM is a tenth of what the pitch says and there is no independent evidence any of it works.
18/30
watch
08
Outbound and GTM-as-a-service
The spread is genuine at 55–75% and the sale closes in weeks, but every retainer is priced against a salary the buyer can look up, and the supply scarcity that made Clay agencies expensive in 2024 was gone by 2026.
18/30
crowded

All 15 verticals · the full scoring table

04

The arguments that cut across every vertical

14 topic pages on the mechanism itself: what a rake can be, why gross revenue is not revenue, what actually stops a buyer and a seller cutting you out, and what the law does to a crowd in December.

05

What this is, and what it is not

189 pages, roughly 305,000 words, and still growing — the atlas was one pass, the niche work is the live edge.

It is

A working record with its seams showing

Every page carries a confidence mark and every number that comes from a leak, a press claim or a company's own blog is marked as such. The open questions page lists what we could not establish, and it is long on purpose.

It is not

Diligence, and not advice

Almost every company here is private and discloses nothing it is not forced to. Take rates are inferred, revenue figures conflict between reputable outlets by factors of two, and the single most important number in the sector — what a frontier lab actually pays for human data — has never been disclosed by anyone.

How this was built and where it is weak · the evidence register