Radiology has the largest device market, the clearest regulatory pull and a genuine, measurable capability gap. It also has an arbitrage that runs backwards, and that single arithmetic fact outweighs everything else on this page.
Radiologists average $571,000 (Becker's/Medscape), implying $286–317/hr of clinical opportunity cost at 1,800–2,000 hours [INFERENCE]. The top rate observed anywhere for medical AI-training work is $300/hr (aigigjobs) [WEAK]. And a productive mixed-modality teleradiologist clears $250–400/hr during active reading, at 80–160 studies per ten-hour shift (natoe.ai) [WEAK — vendor blog].
You are asking the most expensive physicians in medicine to do slower work for less money than the remote piecework they already have set up. Every other constraint on this page is secondary to that one.
What the artefact is
Four products, ascending in value and difficulty:
- Finding-level labels and bounding boxes — the commodity. Priced at $2–5 per chest x-ray for classification, $5–12 with boxes, $40–150 per CT/MRI study for organ segmentation, $100–300 per case for 3D tumour segmentation (mpiricsoftware)
[WEAK — single vendor blog]. - Structured reports — findings mapped to RadLex terminology and RadElement common data elements against RSNA RadReport templates (RSNA). This is what report-generation models are actually scored against.
- Discrepancy adjudication — the double-read plus third-reader pattern the FDA guidance effectively mandates. In one modelled 5,000-study chest CT project, adjudication is a distinct line item of roughly $75K against $350K primary read and $350K second read
[WEAK, same source]. - Reasoning traces on hard cases — why I called this, what I excluded, what I would do next. No observed market price anywhere.
Note that products 1 and 3 are image-native and expensive; product 4 is text and cheap. The labs buy text. See Characterised disagreement.
Does it need patient data
Partially avoidable, and the exceptions are narrower than the field assumes.
MIMIC-CXR gives roughly 377,000 chest radiographs with free-text reports, but it is a research asset, not a commercial input. It is licensed "for the sole purpose of lawful use in scientific research and no other," per individual and non-transferable, with an explicit prohibition on sharing access with third parties — which forecloses the core operation of routing records to a panel of contracted readers. PhysioNet extended that prohibition explicitly to AI services: the DUA "prohibits sharing access to the data with third parties, including sending it through APIs provided by companies like OpenAI, or using it in online platforms like ChatGPT," naming Azure OpenAI with human-review opt-out, Amazon Bedrock, Vertex and Anthropic Claude as acceptable routes (PhysioNet).
CheXpert is non-commercial by default; Stanford's commercial licence is $70,000 per dataset per year (FY25) (Stanford AIMI). TCIA is the one bright spot — most collections are CC-BY 3.0/4.0 with commercial use explicitly permitted, making it the best licence-clean source of CT and MRI (TCIA).
Beyond chest x-ray and TCIA oncology imaging you are into institutional agreements, and that puts an operator in the business of hospital BD rather than clinician recruitment. Two hard walls: in every US state except New Hampshire, the provider or institution — not the patient — owns the medical record (Becker's), so a radiologist cannot bring their own reads. And DICOM carries PHI in the header and burned into pixel text; the ACR warns explicitly that search engines can recover patient identity from image pixels in published material (ACR).
The de-novo route works: a radiologist can compose a clinically faithful case vignette with findings, differential and report, touching no PHI and requiring no pipeline. It is also the lowest-moat version of the product. Full treatment at Build it without ever touching a patient record.
One structural advantage worth naming: de-identified retrospective annotation is not the practice of medicine. A rendered read on an identified patient is, and requires state licensure or the Interstate Medical Licensure Compact. Framing the product as annotation and adjudication rather than opinion sidesteps licensure and creates no physician-patient relationship — the single most important term in the contributor agreement, given that courts have found duty in informal consultations (ASCO Post). See What the doctor on the other end is risking if you need the liability detail.
Is anyone buying
Budget scores 3: strong from one buyer class, near-absent from the other.
Frontier labs: weak. Radiology contributed 8 of 525 examples to HealthBench Professional (PDF). There is not a single frontier-lab job posting anywhere for a board-certified radiologist. Outlier runs a radiology vertical at "up to $120/hr" — but the described tasks are "generating clinical questions," "review and analyze model responses for clinical soundness" and ranking outputs, with no board certification or active licensure required and an explicit note that "clinically experienced NPs, PAs, or nursing professionals may also be considered" (Outlier). That is a cheap text product sold under a radiology label. See The labs as health buyers.
Big Tech research: strong. Microsoft Research fields MAIRA-2 and CARE-X on ReXrank. MedGemma 1.5 trains on MIMIC-CXR, ChestX-ray14, CheXpert, MS-CXR-T and Chest ImaGenome plus private de-identified CT and MRI datasets from a US diagnostic centre, evaluated against "radiologist adjudicated labels" (model card). Someone is being paid for those private sets and those adjudicated labels; nobody has said who.
Device sponsors: strongest. There were 1,016 FDA authorisations through 20 December 2024 representing 736 unique devices, and 88.2% of image-based devices are radiology (npj Digital Medicine). The same paper notes it "did not find evidence of large language models in the studied device list" — the regulated device market and the frontier-LLM market are, on that analysis, disjoint buyer populations. The January 2025 draft guidance requires documented annotator expertise, blinding and adjudication, which a sponsor cannot satisfy with anonymous crowd labels: The regulator wrote your product spec.
Proof scores 3. ReXrank is open and contested, so a well-built reader study is legible and citable — but the audience that would cite it is academic, and the audience with money is a device sponsor whose procurement runs through an institutional data agreement you do not control.
What the expert costs
Cost scores 1, the lowest in the dossier, and it is the whole verdict.
Opportunity cost is $286–317/hr. Observed AI-data rates run $72/hr (OpenTrain evaluator), $150/hr (OpenTrain board-certified premium tier) and $300/hr (Handshake) [WEAK]. A cross-check puts median radiology production at 10,500 wRVU/yr against $470K median compensation at $42–48/wRVU, and models a teleradiologist at ~$550K on a 40-hour week, about $264/hr (fastrvu) [WEAK — planning source].
Add the pilot costs the other niches do not carry: DICOM de-identification, a viewer, storage, and either a $70k/dataset licence or a hospital agreement. A niche needing licences and infrastructure before the first read scores 1 on cost, and this one does. Client-side comparators for context: teleradiology list prices run $12 x-ray, $28 ultrasound, $32 mammography, $40 CT, $60 MRI, $99 PET-CT per study (NDI); Braid Health sells a consumer second opinion at $199 per read (Radiology Business). Compare those to $286–317/hr of physician time and the margin structure is visible.
Getting to them
Reach scores 4. The ACR has more than 39,000 members (ACR); RSNA drew roughly 38,000 attendees in 2025 (AuntMinnie); teleradiology platforms are the highest-yield channel because those radiologists are already remote, already 1099 and already tooled for piecework; and NPI taxonomy filters give a clean national frame.
Workforce dynamics cut both ways. Radiologist numbers grew 17.3% between 2014 and 2023 while practices consolidated (9.7 to 17.9 radiologists per practice), subspecialty radiologists are 37% more likely to exit the workforce, and radiologists are leaving practice entirely at over twice the rate of a decade ago (ACR Bulletin 2026). Scarcity is good for defensibility and bad for recruiting cost — see What a clinician hour costs.
Reach is 4 rather than 5 because the channels are excellent but the pitch is weak: you are competing for hours against a piecework market that pays more.
Where the benchmarks sit
| Benchmark | Status | Best score |
|---|---|---|
| ReXrank (4 datasets incl. private ReXGradient, 10,000 studies) | Open, actively contested | Deepwise-RG and Microsoft CARE-X lead; MAIRA-2 third on ReXGradient |
| ReXVQA | Open | CARE-X 90.83%, CheXOne-R1 88.03%, MedGemma-4B-it 83.44% |
| CXR report generation overall | Explicitly unsolved | PSB 2026 title: "Automated Chest X-ray Report Generation Remains Unsolved" |
| RadBench (Harrison.ai) | Open framework | scores not published [UNVERIFIED] |
| MedGemma 1.5 internal | — | CT 61.1%, MRI 64.7%, MIMIC-CXR F1 89.5% |
| GPT-family on radiology cases (28 studies, 8,852 cases) | Far from saturated | GPT-4T 72.0%, GPT-4o 57.2%, GPT-4 56.5%, GPT-4V 42.3% |
Sources: ReXrank, PSB 2026, RadBench, MedGemma, Frontiers in Radiology. See The measurement landscape.
The headline is the last row: GPT-4V at 42.3% is materially worse than text-only GPT-4 at 56.5% on the same radiology cases. Frontier vision is worse at radiology than frontier text is at radiology case descriptions. That is a large, open, image-native capability gap — precisely the one image-native expert data would close, and precisely the one nobody has been shown to be funding.
What would kill it
Defense scores 3 and room scores 3. Supply stays scarce and imaging keeps changing, so the data does need refreshing — but the position is contestable from three directions.
The economics kill it first. There is no wage arbitrage, and margin would have to come entirely from the credentialing and adjudication wrapper rather than from the labour spread.
The ACR is the second threat, and it is not hostile so much as proprietary. It launched Assess-AI, "the world's first AI quality registry," on 18 November 2024, collecting AI concordance with radiologist reports across real deployments (ACR). ACR is positioning itself as the neutral arbiter of imaging-AI performance — which is the discrepancy-adjudication product, run by the society that owns the members.
Third, institutional data ownership means the sponsor's data agreement, not your roster, is the scarce asset.
There is no disclosed contract between any frontier lab and any specialist-physician data vendor. The single strongest radiology data point — Handshake advertising radiologists at up to $300/hr for work "supporting AI research for companies like OpenAI and Anthropic" — is one aggregator blog reporting a job listing, with no lab named as counterparty and no deal size attached.
Where the record is thin
The definitive source on radiology discrepancy rates, "Forty-One Million RADPEER Reviews Later" in JACR, is robots-blocked [UNVERIFIED]. Without those base rates the discrepancy-adjudication product cannot be sized — pathology has its equivalent number (53.5% unanimous agreement), radiology does not. The per-study and per-image annotation prices rest on a single vendor blog. And nobody has identified the "US diagnostic center" supplying MedGemma's private CT and MRI, which is the one confirmed radiology data purchase in the record.
Ranked comparison at The health read; the commodity backdrop at Expert data for frontier labs; the incumbent at Centaur.ai, in full.