Ophthalmology is the only specialty in medicine where an autonomous AI diagnosis is a billable service. CMS finalised a national payment rate for CPT 92229 — autonomous point-of-care retinal imaging with AI — effective CY2022, three years after LumineticsCore's 2018 De Novo authorisation (Digital Diagnostics). The AAO has since had to advocate against a Medicare carrier under-pricing the code (AAO).
That single fact reorganises the analysis. Reimbursement means device sponsors here have revenue rather than runway, and a sponsor with revenue can pay recurrently for reference-standard grading. No other specialty in this dossier has a buyer whose product is already on a fee schedule.
The second fact is the product thesis. BELO, an ophthalmology benchmark of 900 questions, reports OpenAI o1 at accuracy 0.882 and macro-F1 0.890 — but text-generation quality scores of only 20.40–71.80 out of 100 (arXiv 2507.15717). The answers are nearly right. The reasoning is not. That gap is what to sell.
What the artefact is
Dermatology: lesion diagnosis labels on clinical and dermoscopic photography, ideally biopsy-confirmed; segmentation masks; Fitzpatrick-stratified labels, where the fairness axis is a first-class product in a way it is nowhere else; triage decisions — biopsy, reassure or refer; and rubric-scored diagnostic narratives.
Ophthalmology: diabetic-retinopathy severity grading on fundus photographs against the ETDRS scale — the single most commoditised medical-image label in existence; OCT layer segmentation and pathology labels; glaucoma disc and field interpretation; referral-threshold decisions; and written clinical reasoning.
BELO is the most useful single artefact in the whole dossier because it itemises the expert labour: 900 MCQs drawn from BCSC, MedMCQA, MedQA, PubMedQA and BioASQ, curated by roughly 23 ophthalmology professionals across three tiers — one board-certified ophthalmologist plus two optometrists plus six research staff for QC, 10 board-certified ophthalmologists refining explanations, 3 senior ophthalmologists adjudicating (Ophthalmology Science). That is a costable, repeatable production recipe with a named tier structure. Copy it.
On the dermatology side, DermBench built 4,000 images with expert-certified diagnostic narratives and rated 4,500 model narratives across six dimensions — accuracy, safety, medical groundedness, clinical coverage, reasoning coherence, description precision — with best accuracy around 3.2 out of 5 (arXiv 2511.09195). Six-dimensional narrative rating is the sellable artefact, not the diagnosis label.
Does it need patient data
Less than any other imaging specialty, but the permissiveness is routinely overstated and the detail matters.
ISIC is the largest public dermoscopy archive and it is licensed per image, not per archive. Counts retrieved from the ISIC API on 30 August 2026: CC-0 48,751; CC-BY 258,980; CC-BY-NC 245,288; total 553,019. Roughly 44% of ISIC is off-limits to a commercial product, and the split is invisible unless you check per-image metadata — meaning anyone who bulk-downloaded "ISIC" and trained on it has probably breached CC-BY-NC without noticing. The ~307k CC-0/CC-BY subset is a legitimate commercial substrate; the rest is not. Two consequences: filter per image, and note that licence-hygiene auditing is itself a saleable service, because buyers almost certainly have this contamination already. Full treatment in Build it without ever touching a patient record.
HAM10000 is published in Scientific Data (Nature). MRA-MIDAS (Stanford AIMI) pairs clinical and dermoscopic photography with histopathologic ground truth from board-certified dermatopathologists, with consensus review for severe dysplasia and independent dermatologist verification — the highest-quality public derm reference standard (Stanford AIMI). But Stanford AIMI's commercial route is a paid enterprise licence: online application, committee review, and $70,000 per dataset per year (FY25) (Stanford AIMI). "Permissive" here means "priced".
EyePACS is the canonical public diabetic-retinopathy fundus corpus and sits in MedGemma's training mix — but its actual licence terms could not be verified to primary-source standard [UNVERIFIED]. It is best known through a Kaggle competition, and competition rules typically grant a licence for participation and academic research rather than unrestricted commercial use. Do not treat "it was on Kaggle" as evidence of a permissive licence.
The de-novo route works well here and has a feature unique to dermatology: a skin photograph taken by the patient on a phone is not institutionally owned. A derm-focused operator can commission consented image collection directly, which is impossible in radiology and pathology. And annotation of de-identified images is not the practice of medicine, so no licensure or malpractice duty attaches — the boundary must be explicit in the contributor agreement.
Is anyone buying
Budget scores 4. The evidence is good, plural and still inferential.
- Handshake AI advertises ophthalmologists at up to $300/hr, jointly the top medical rate observed anywhere (aigigjobs)
[WEAK — single aggregator, 95 listings sampled]. - Dermatology is a top-represented specialty in both HealthBench and HealthBench Professional; ophthalmology contributed 18 of 525 HealthBench Professional examples — small in absolute terms, but 6× radiology's 8 and 6× pathology's 3 (PDF).
- Google is the deepest buyer. MedGemma 1.5 (13 Jan 2026) required licensing six private de-identified dermatology datasets from Colombia, Australia and Japan plus an internal Fitzpatrick 5–6 collection, alongside EyePACS, and reports EyePACS accuracy 76.8% (model card). Someone was paid for those six collections.
- DermEVAL was published at WACV 2026 as "a dermatologist-reviewed benchmark" (CVF) — dermatologist-review-as-a-service has at minimum a research market.
- The device pull is the strongest of the six. See The regulator wrote your product spec: the FDA's January 2025 draft guidance asks sponsors to document annotator expertise, blinding, adjudication methodology and inter-clinician variability as named submission content, plus demographic representativeness — which is where Fitzpatrick-balanced labelling stops being a virtue signal and becomes a regulatory deliverable.
Proof scores 4. A BELO-shaped replication in dermatology, or a reasoning-quality rating set on the DermBench dimensions, is publishable within a quarter and would make an operator the visible best source. The population is small enough that ten senior names are a credible panel.
What the expert costs
Cost scores 3, and the average understates the difficulty.
Dermatology averages $448K and ophthalmology $446K, implying $223–249/hr at 1,800–2,000 hours [INFERENCE] (Becker's/Medscape). The observed $300/hr for ophthalmologists sits above that average — the only case in the dossier where the AI-data rate beats clinical opportunity cost.
But both specialties are heavily procedural: Mohs, cosmetics, cataract, injections. Their marginal hour is worth considerably more than their average hour, so the real arbitrage is worse than $300 versus $223 suggests, and recruiting will select for the non-procedural minority — medical dermatology, comprehensive ophthalmology, part-time and retired. Budget for image licensing (Stanford at $70k/dataset/year), per-image licence filtering and a viewer, on top of labour.
Getting to them
Reach scores 5, and it is earned on one number: AAO's 32,000 US members represent more than 90% of practising US ophthalmologists, plus over 7,000 international members (Wikipedia) [WEAK — not primary]. A list covering 90% of a specialty is a census, not a channel. AAD carries >21,000 members (Wikipedia) [WEAK].
Beyond the societies: AAO's IRIS Registry, the largest specialty clinical registry in medicine; ISIC's own collaborating-investigator network, which is a ready-made dermatologist annotation community and the single highest-yield channel here; retina and glaucoma subspecialty societies; teledermatology platforms, structurally analogous to teleradiology in that the clinicians are already remote and already 1099; and NPI taxonomy filters. See What a clinician hour costs.
Society posture is the friendliest of the imaging specialties — the AAO is lobbying for better AI reimbursement rather than resisting AI.
Where the benchmarks sit
| Benchmark | Status | Score |
|---|---|---|
| BELO (900 MCQs, ~23 curators) | Open; expert cost documented | o1 accuracy 0.882; text quality 20.40–71.80/100 |
| Ophthalmology board-style questions | Near-saturated | o1 beats GPT-4o, Gemini 1.5 Flash and human test-takers |
| DermBench (4,000 images, 4,500 rated narratives) | Open, wide gap | best accuracy ~3.2/5; worst 1.2 (LLaVA-Med-7B) |
| DermEVAL | New, WACV 2026 | dermatologist-reviewed |
| MedGemma 1.5 on EyePACS | — | 76.8% |
Sources: BELO, Ophthalmology Science, DermBench, MedGemma model card. Wider context in The measurement landscape.
The BELO split — accuracy 0.882, reasoning 20–72 — is the whole thesis in one row.
What would kill it
Defense scores 4 and room scores 4, but two things could close it.
The population is small. 21,000 dermatologists and 32,000 ophthalmologists is a feature for a boutique and a ceiling for anything larger; if a competitor signs the fifty best-known names first, there is no second tier of equivalent authority. Centaur.ai, in full shows what a funded medical-image labelling incumbent looks like when it gets there first.
And the reference standard is unreliable in exactly the place that matters most. Eight board-certified dermatopathologists reviewing 792 melanocytic cases agreed unanimously in only 53.5%, with κ = 0.428 for non-invasive melanoma (Nature Communications). Sell melanoma labels as facts and the product is falsifiable; sell them as an adjudicated distribution and it is not. See Pathology.
There is no disclosed contract between any frontier lab and any specialist-physician data vendor, on this sub-niche or any other on the diagnostic side. The $300/hr ophthalmology rate rests on one aggregator blog. Google's six private dermatology collections are named in a model card with no supplier, price or term. Every buying claim on this page is inference from rate cards, model cards and benchmark acknowledgements.
Where the record is thin
The actual dollar value of CPT 92229 is not disclosed in the source announcing it [UNVERIFIED] — which matters, because the size of the reimbursement determines how much a sponsor can pay for grading. AAD and AAO annual-meeting attendance figures were not published on accessible pages, so channel sizing is soft. And the EyePACS licence remains unverified; nothing should be built on it until the terms are read.
Ranked comparison at The health read; the strategic read at Characterised disagreement.