Miju Labs

All companies

Surge AI

Bootstrapped to over $1B of revenue with 130 employees and no outside capital — and the one company whose valuation nobody can pin down.

medium confidence5 minupdated 2026-08-29ai labs · data · bootstrapped
Vertical
Expert data for frontier labs
Founded
2020 (Sacra says 2021)
Headquarters
San Francisco
Raised
$0 until 2025; sought up to $1B in July 2025 — no confirmed close
Last valuation
Reported $15B (Reuters) or $25B (Bloomberg) — sources conflict and no close is confirmed
Revenue
~$1.2B annualised 2024, basis not stated but almost certainly GROSS. No credible 2025 or 2026 figure.
Status
Active, profitable since launch; Edwin Chen reportedly owns ~75%
Who runs it · 2 people in the index

Latest (Sep 2026): Contrary Research (Aug 2026) says no outside round has closed despite the July 2025 talks at a reported $15B-$25B+ valuation, so Surge is still bootstrapped and about 75% owned by Edwin Chen. In 2026 Surge has repositioned as a research-led lab: Hemingway-bench, a writing-quality leaderboard judged by expert writers (Feb 4), GDP.pdf (Apr 14), cited in Anthropic's and OpenAI's 2026 model cards (Jun-Jul), and the Tuesday Work Index (Aug 18). A Sep 16 2026 post covers helping Anthropic build automated alignment researchers. No verified 2025 or 2026 revenue figure; the last credible one is $1.2B for 2024.

What Surge AI is saying
one of my favorite examples from our Kimi coding post-training run: it had to write a Zstandard decompressor, but there was no zstd binary available to check whether it worked. so the trained model wrote a compressor first, generated its own valid test files, and used those to test the decompressor. no ground truth existed, so it built one itself. Read the full report here: surgehq.ai/blog/hill-climbing-sw…
Edwin Chen
Founder, Surge AI
one of my favorite examples from our Kimi coding post-training run: it had to write a Zstandard decompressor, but there was no zstd binary available to check whether it worked. so the trained model wrote a compressor first, generated its own valid test files, and used those to test the decompressor. no ground truth existed, so it built one itself. Read the full report: lnkd.in/e3pv9gnR
701 comments2 reposts
Surge AI@HelloSurgeAI·
Anthropic recently published new work on automated alignment researchers: Claude agents that search the literature, propose alignment methods, train models, evaluate the results, and iterate. Across ten alignment failures—including deception, sycophancy, jailbreaks, privacy violations, and reward hacking—the automated researchers found methods that improved safety benchmarks while preserving general capabilities. The strongest methods also generalized to held-out benchmarks, open-ended Petri audits, and models up to 4.7× larger than the models they optimized against. We contributed to Anthropic's research by building and running the human researcher baseline. Anthropic compared its automated researchers with ideas from 28 experienced technical AI safety researchers, each given up to eight hours to propose a method for addressing the same alignment failures. Surge ran that pipeline end to end, including researcher recruitment, structured submissions, quality control, and expert review. As the paper puts it: “The human baseline is collected with Surge AI, whose pipeline the study runs through end to end.” The automated researchers ultimately found methods that outperformed the human-proposed baselines on the seven alignment failures where human ideas were collected. Anthropic is careful about the comparison—the agents could iterate repeatedly, while the human researchers submitted one idea—but the work offers a compelling glimpse of how automated research might complement human researchers in the future. We spend a lot of time at Surge thinking about what happens as models take on increasingly expert work: how to build credible human baselines, how to evaluate work that requires real judgment, and how to turn expert human knowledge into useful training and evaluation signals. This study continues our research collaboration with Anthropic that goes back to training Claude with expert human feedback, and research on scalable oversight, inverse scaling laws, and Constitutional AI. We’re glad to have played a small part in this one. Read Anthropic’s research⁠: anthropic.com/research/automated…

Surge AI is the counterexample that the rest of the expert data vertical would rather not discuss. Founded by Edwin Chen, an ex-Google and ex-Meta engineer, it reached roughly $1.2B of annualised revenue in 2024 — bootstrapped, profitable since launch, with about 130 full-time employees. No source states whether that figure is gross or net; on every comparable in this vertical it would be gross, and the whole margin question below turns on it (Sacra). That is ~$9M of billings per employee, and it beat Scale AI's $870M in the same year on $1.6B less capital.

What it does differently

Four things, all reported rather than claimed:

  1. Quality as the wedge. Chen's founding observation, from inside Google, Twitter and Facebook, was that vendor data was full of mislabellings done for minimal pay by people without relevant backgrounds.
  2. Paid workers more on purpose. Sacra puts contractor pay at "30–40 cents per working minute" — $18–24/hour (Sacra).
  3. Refused to build a sales org. 130 FTEs against ~50,000 expert contractors.
  4. Refused capital, which preserved neutrality — precisely the asset Scale AI destroyed the moment Meta took 49%. When Meta's relationship with Scale soured, Meta moved work to Surge and Mercor (TechCrunch).

Customers: roughly 12 frontier labs generating over $1B of revenue, with OpenAI, Google, Anthropic, Microsoft and Meta named (Sacra; Wikipedia). Chen has said Surge works with clients paying eight- or nine-figure contracts (Inc via Yahoo) — the only concrete contract-size disclosure found anywhere in this vertical.

The valuation nobody can confirm

Gap in the record

In July 2025 Surge hired advisers to raise up to $1B. Reuters reported a $15B+ valuation. Bloomberg reported "at least $25B." A LinkedIn post claims $1B raised at $30B. No primary source or tier-one publication confirms that any round ever closed, at any valuation. This is a $10B ambiguity sitting in the middle of the sector's league table, and it should not be resolved by picking the number you prefer.

Sources: Sacra, Bloomberg, SiliconANGLE, LinkedIn (unverified).

Revenue is barely firmer. The last credible figure is $1.2B for 2024. Latka claims $1.4B for 2025, self-described as an estimate and tagged [WEAK] in the research; Latka's own Surge page has failed spot-checks badly enough that its figures should not be used as evidence at all.

The margin question

Surge does not disclose a take rate or gross margin. Sacra says only "strong gross margins."

Arithmetically, $1.2B of revenue on ~130 staff and $18–24/hour labour implies a margin materially above Mercor's leaked 27–33%. But that inference depends entirely on whether Surge's $1.2B is gross or net, and no source says which. If it books gross like everyone else in this vertical, the implied margin is high but unbounded from below by any evidence. See GMV is not revenue and What a rake can actually be.

Gap in the record

Surge's gross margin, take rate, 2025 revenue and 2026 revenue are all unknown. What is on the record is that it was profitable from launch and that Chen reportedly still owns about 75% of it.

Trouble

  • Misclassification class action, 21 May 2025 (named plaintiff Dominique DonJuan Cavalier II, Clarkson law firm), alleging deliberate misclassification of annotators as contractors, unpaid training, and wage deductions arising from impossible task time limits (AOL/SF Standard). See The law is about to arrive.
  • July 2025 leak of internal RLHF guidelines used for Anthropic training (Wikipedia).
  • China. Surge is reported to have actively courted Chinese labs (AI Weekly summarising Forbes) [WEAK — second-hand]; Alexandr Wang publicly attacked Surge and Mercor over it (Wang on X).

DataAnnotation.tech

Widely believed to be Surge's consumer-facing recruiting front. Wikipedia notes criticism of its "lack of ownership transparency."

Not verified

The ownership link between Surge and DataAnnotation.tech is not confirmed by any authoritative source. Treat it as unverified.

The platform advertises $20–30+/hr for generalist and multilingual work, up to $60/hr for coding and $50–100+/hr for STEM and professional tasks (DataAnnotation). Independent trackers put the realistic average nearer $20/hr, with churn and unexplained account deactivations the dominant complaint theme — both [WEAK], from community-data blogs rather than press.

The read

Surge is the strongest evidence in the atlas that this business does not need venture capital, and the strongest evidence that neutrality is worth more than a strategic investor's cheque. It is also the least legible company in the vertical: one hard revenue figure, two years stale, and a valuation with a $10B error bar.