Miju Labs

All companies

Appen

The only pure-play with audited numbers, and a 97% drawdown from peak — the base rate for what a concentrated data vendor is worth when one hyperscaler leaves.

high confidence5 minupdated 2026-08-29ai labs · data · public comps · concentration
Vertical
Expert data for frontier labs
Founded
1996; ASX-listed since 2015
Headquarters
Sydney
Raised
Public company; an A$60M placement in 2024 is unverified
Last valuation
A$330M market cap (28 August 2026) — roughly 0.9x revenue, computed AUD-on-AUD against A$361.9M TTM. Peak was ~US$4.3B in August 2020
Revenue
US$230.8M FY2025 operating revenue (audited), 40.3% gross margin, US$12.2M underlying EBITDA. NET — Appen employs or contracts its crowd and books the work, not a payout marketplace
Status
Survivor, shrunken. Loss-making every year since 2022; stabilised and growing again on an AI-data tailwind
Who runs it · 6 people in the index

Latest (Sep 2026): FY2025 results (25 Feb 2026): operating revenue of US$230.8M, underlying EBITDA of US$12.2M (up 251%), Q4 gross margin of 45%, and Appen China up 75% to US$102.9M. Appen also gave FY2026 revenue guidance of US$270-300M. Vanessa Liu became Non-Executive Chair on 1 Jan 2026, replacing Richard Freudenstein, and a new General Counsel (Jaime Frasca) started in Jan 2026. At the 22 May 2026 AGM, shareholders approved CEO Ryan Kolln's incentive grants. Market cap was ~A$330M (28 Aug 2026), against a ~US$4.3B peak in 2020.

What Appen is saying
Appen
1,076,785 followers
We’re publishing MedTerm-90, a benchmark of seven speech-to-text systems on 90.5 hours of real physician dictation. The benchmark evaluates 1,413 recordings across 14 specialties and scores 30,017 clinician-validated entities using clinical entity recall rather than word error rate. The systems tested: • ElevenLabs Scribe v2 • OpenAI gpt-transcribe • Microsoft AI MAI-Transcribe 2 • Google Gemini 3.5 Transcribe • Meta Muse Transcribe • Amazon Transcribe Medical • XAI Grok STT 1.0 ElevenLabs Scribe v2 ranked first at 83.35% recall, followed by OpenAI gpt-transcribe at 81.79% and Microsoft MAI-Transcribe 2 at 81.31%. But the aggregate ranking is only part of the result. On 28.0% of recordings, no system reached 85% clinical entity recall. We also reviewed 6,425 near-miss renderings for clinical consequence. 191 were classified as potentially fatal and 1,027 as serious. The errors were often plausible transcription outputs rather than obvious failures. Examples included drug substitutions, incorrect dosages, and clinically opposite terminology such as hyponatremia rendered as hypernatremia. The evaluation uses real, previously unseen physician dictation rather than synthetic or studio-read speech. Seventy percent of the corpus is GSM-compressed telephony audio, with the remainder coming from MP3 and uncompressed WAV recordings used in real transcription workflows. We also manually reviewed 12,326 distinct near-miss pairs. Only 21.2% were accepted as clinically equivalent written forms. Full results include confidence intervals, speed measurements, category-level performance, methodology, judge policy, risk analysis and model-specific failure patterns. Benchmark, methodology and error analysis in the comments. #ASR #SpeechRecognition #MedicalAI #AIResearch #Benchmarking #SpeechToText
333 comments7 reposts
Ryan Kolln reposted this
Appen
1,076,785 followers
VLM judges can be right about the details and still get the overall judgment wrong. A judge is only working with the frames or clips it’s given, so if the important moment never gets retrieved, better reasoning downstream won’t fix it. And even when it correctly catches a missing detail, there’s still another question: did that omission actually matter to the quality of the output? That gap between detecting an error and deciding how much it should affect the score is a big part of VLM evaluation. The research also shows that temporal reasoning remains hard even on short videos, which means this isn’t just a context-window problem. Better evaluation needs better retrieval, better rubrics, clearer weighting of what matters, and human review where the judgment is ambiguous. Jeanine Sinanan-Singh, Director of GenAI Research at Appen, digs into this in our latest blog, including what a stronger multi-resolution evaluation stack can look like. Link in the comments.
316 comments4 reposts
Appen
1,076,785 followers
Healthcare AI may be approaching a bottleneck that better architectures and more compute won’t solve: access to the right data. Clinical AI is moving beyond narrow benchmarks into multimodal systems that need to reason across imaging, clinical notes, biosignals and physician-patient dialogue. But the data required to train and evaluate those systems is unusually difficult to scale. It has to be privacy-safe and traceable. It often requires credentialed medical experts to annotate and adjudicate. And for many emerging use cases, public or synthetic datasets don’t capture the clinical variability models encounter in deployment. The problem becomes even sharper with clinical agents. Recent research shows how dramatically performance can change when models move from static medical questions to sequential, tool-using doctor-patient interactions. That means evaluation itself needs richer clinical data: longitudinal dialogue, multiple specialties and languages, realistic tool use, and expert validation. As healthcare models become more capable, the limiting resource may increasingly be high-quality, expert-validated clinical data rather than model capacity. Our latest Appen Insights looks at the coming medical data crunch, and why data access, provenance and clinical expertise could become a competitive moat for healthcare AI. Read the full blog by Sergio Bruccoleri, VP of Delivery at Appen, in the comments. #HealthcareAI #AIResearch
243 comments1 reposts

Appen is the base rate. It is the only pure-play in the expert data vertical that has run a full cycle in public with audited numbers, and everything the private cohort is currently being marked on has already been tested here.

The cycle

YearRevenue (US$)Net income (US$)
FY2021$447.3M$28.5M
FY2022$388.3M–$239.1M
FY2023$273.8M–$118.1M
FY2024$235.2M–$20.0M
FY2025$232.7M–$21.8M

Source: stockanalysis.com. A 48% revenue decline over four years, with the equity value falling further: peak market capitalisation surpassed the equivalent of US$4.3 billion in August 2020, against A$330M on 28 August 2026 — a drawdown The Verge computes at 97%. (The atlas's own two figures do not quite reproduce that. A$330M against the ~US$4.3B peak converted to Australian dollars gives about 92%; converting the current cap to US dollars instead and comparing to US$4.3B gives about 95%. The 97% is carried here as reported by The Verge, not as computed by this page — the gap is a currency-basis artefact of exactly the sort item 7 of the atlas's contradiction list warns about, and it does not change the finding, which is that essentially all of the equity value went.)

Appen's collapse was not caused by bad execution on labelling. It was caused by buyer concentration plus a change in training technique. Both conditions are present, in more extreme form, across the current private cohort. See One customer is a binary event.

The Google contract

At peak, 80% of revenue came from five clients — Microsoft, Apple, Meta, Google and Amazon. On 22 January 2024, Alphabet terminated a contract worth roughly US$83M, about a third of remaining revenue, as it cut thousands of search quality raters (NBC). Appen shares fell 40–41% in a single day. North American offices closed the following month and executives left through May 2024.

One buyer decision, one third of the revenue, one day. That is the mechanism the entire vertical is exposed to, and it is why Mercor's ~91% revenue share from foundation-model companies is the most important number on its page.

FY2025 — stabilisation, and what it cost

The most recent audited year is more interesting than the headline suggests (Appen FY2025 Annual Report):

  • Operating revenue US$230.8M, underlying EBITDA US$12.2M — up 250% from $3.5M.
  • Gross margin 40.3%. Higher than Mercor's leaked 33%.
  • Appen Global fell 21.1% to $127.9M. Appen China grew 74.8% to $102.9M. The business is now nearly half Chinese.
  • Generative-AI revenue rose to 33% of total, from 22%.
  • Top five customers = 74.3% of revenue, up from 67.3%. Concentration is increasing, not falling.
  • Headcount 1,185; crowd of 1M+ contributors across 200+ countries and 500+ languages.
  • Crowd NPS fell from 33 to 22.

That last line deserves more attention than it gets. Worker satisfaction deteriorated even as the business stabilised — the same pattern visible in falling rates at Outlier, Mercor's Musen-to-Nova cut and Handshake's Project HH. The supply pools are shared across platforms, so a deteriorating crowd is a real cost of goods, not a PR problem.

What the market pays

~0.9x revenue — A$330M of market cap against A$361.9M of TTM revenue, both in AUD — on a 40.3% gross margin business that is growing again. Innodata, the other public comparable, trades at ~6.2x on a ~40% gross margin (49% adjusted, Q2 2026) and 40%+ growth. Both sit far below the private marks: Mercor at roughly 29–38x net, Invisible Technologies at ~15x, Handshake AI at ~7.8x net.

So what

Appen at 0.9x and Innodata at 6.2x are not opinions. They are the market's price for this business when it can see inside it. Everything above 6x in the private cohort is a bet that those companies are structurally different from the public ones — and the difference has to come from growth rate and gross margin, because the labour network is demonstrably not what gets paid for. iMerit, a decade-old annotation company with an expert network, sold to EXL for up to $310M in the same quarter Mercor was talking at $20B.

See What the public market pays for labour for the full restatement.

Gap in the record

Two revenue figures circulate for FY2025 — US$230.8M in the annual report's operating-revenue line and US$232.7M in the market-data record — and a TTM figure of A$361.9M (+11.5%) is quoted on a different currency basis. The differences are small but they are not reconciled, and the AUD and USD series should never be mixed.

Not verified

The A$60M placement attributed to 2024 is unverified. So is any claim about Appen's current customer names beyond the audited concentration disclosure.