Most vendors in this market are guessing at what a buyer wants their annotation process to look like. In medical devices you do not have to guess. FDA has published the specification, in a numbered list, in a draft guidance that is still open for comment.
"Artificial Intelligence-Enabled Device Software Functions: Lifecycle Management and Marketing Submission Recommendations," draft guidance for industry and FDA staff, issued 7 January 2025, docket FDA-2024-D-4488 (landing page, full PDF). Still draft as of 30 August 2026, comment period still open — which matters, because a draft is a statement of what reviewers already expect to see, not a future obligation.
FDA explains why it cares about your data at all in a sentence worth memorising:
"For an AI-enabled device, the model is part of the mechanism of action. Therefore, a clear explanation of the data management, including data management practices (i.e., how data has been or will be collected, processed, annotated, stored, controlled, and used) and characterization of data used in the development and validation of the AI-enabled device is critical for FDA to understand how the device was developed and validated."
If the model is the mechanism of action, the data is not an input to the device. It is part of the device, and it inherits device-grade documentation obligations. That single move is what converts clinician annotation from a cost line into a regulatory deliverable.
What the guidance actually asks for
Section VIII on Data Annotation is the shortest and most commercially loaded passage in the document:
"A description of the expertise of those performing the data annotation."
"A description of the specific training, instructions or guidelines provided to data annotators to guide their annotation decisions, including whether annotators are blinded to each other."
"A description of the methods for evaluating quality/consistency of data annotations and adjudicating disagreements (consensus evaluation, sampling). FDA recommends the use of independent assessments by each annotator, without knowledge of the other annotators' decisions, to ensure objective high-quality data annotations"
"A detailed plan for addressing incorrect data annotation."
The parallel Reference Standard subsection adds, where the standard is based on clinician evaluations: the grading protocol used; what data are provided to those clinicians; the "blinding protocol" and "number of participating clinicians and their qualifications"; and
"An assessment of the intra- and/or inter-clinician variability for each task, as applicable, as well as an assessment on whether the observed variability is within commonly accepted standards for a particular measurement task."
Read as a product specification, FDA is asking every AI device sponsor to produce, per dataset: named annotators with documented credentials; a written grading protocol and the instructions actually issued; independent, mutually blinded assessments; a quantified inter- and intra-rater analysis; a benchmark judgement on whether that variability sits within accepted standards; a documented adjudication procedure; a remediation plan for incorrect annotations; and an explicit statement of the uncertainty in the reference standard itself.
Almost no AI company can produce that retrospectively from an internal or crowd-labelling process.
Item four cannot be reconstructed. That is the moat.
Seven of those eight items can, with enough effort and embarrassment, be written up after the fact. Credentials can be collected late. A grading protocol can be reverse-engineered from what annotators were told. A remediation plan is a document.
The inter-rater variability assessment cannot. If annotators were not independent and mutually blinded at the moment of collection, the statistic does not exist. It is not missing paperwork; it is a measurement that was never taken, about an event that has passed. No budget spent in 2027 creates an inter-rater kappa for labels collected in 2025 by three people in a shared spreadsheet.
This is the whole defensible position. A vendor that ships labels with a conformant reference-standard dossier is selling something structurally different from a labelling shop — not better labels, but labels whose production is provable. And because the proof has to be manufactured concurrently, a buyer who skipped it has only two options: re-annotate from scratch, or go into the submission with a hole. Both are your revenue.
Three consequences follow for how the business is run, and they are non-obvious.
Build the blinded, three-reader pipeline before the first customer, not after. The pipeline is the product. A vendor who runs a cheap single-reader process for early customers and plans to "upgrade later" has produced a corpus that can never be retrofitted, and every item in it is permanently second-tier.
Never destroy the individual assessments. Collapsing a panel to a consensus label throws away the exact artefact FDA asks for by name. Store every independent read, blinded, with annotator identity and credentials, indefinitely.
Price on the dossier, not the label count. A buyer comparing your per-item price against a crowd platform's is comparing the wrong things, and the answer to that objection is the deficiency letter they will otherwise receive. See What actually gets sold for where the sellable unit sits.
The reframing that matters most: FDA has already conceded the expert is not truth
The traditional structure of medical AI validation runs: expert equals ground truth, model is scored against expert, model accuracy is bounded above by expert accuracy. When models start matching or beating readers, that structure breaks in a specific way — you can no longer distinguish "the model is wrong" from "the reference standard is wrong."
The guidance shows the regulator has already stopped assuming the expert is right. Among the required descriptions of the reference standard:
"A description of the uncertainty inherent in the selected reference standard."
"A description of the strategy for addressing cases where results obtained using a reference standard may be equivocal or missing."
FDA is not asking sponsors how accurate their experts are. It is asking them to quantify how far from truth the expert is — and to have a policy for the cases where truth is unavailable.
A single expert opinion has no error structure. It is a point with no variance, and there is nothing to write in the "uncertainty inherent in the reference standard" section except a shrug. A panel of three or more, independently blinded, with measured agreement and a documented adjudication log, produces a label and its error bar. Only the second thing is submittable.
This is why the product appreciates as models improve rather than depreciating. As a model approaches single-reader performance, the value of one more single-reader label falls toward zero — the model already produces those. What rises in value is the structure around the disagreement: where do competent readers diverge, by how much, and how does the divergence resolve. That region is where the residual clinical uncertainty lives, and it is exactly where a device needs a confidence threshold or a human in the loop.
It is also why you should refuse to sell single-annotator labels even when a buyer asks for them and they would be cheaper to produce. It builds the wrong asset, it cannot be upgraded, and it trains the market to price you as a labelling shop. This is the one commercial rule in the dossier that should be treated as absolute. Pathology and Radiology are where the discipline pays hardest, because both are domains with well-documented reader variability that a panel can actually measure.
The recurring-revenue mechanism most people miss
The other guidance to read is the Predetermined Change Control Plan final guidance, docket FDA-2022-D-2628 (FDA) [WEAK] on the publication date — the FDA page reports August 2025 while the guidance was finalised in late 2024, so verify before citing a date to a client.
A PCCP has three parts: the Description of Modifications, the Modification Protocol, and the Impact Assessment. Its purpose is to let sponsors implement pre-authorised changes "without necessitating additional marketing submissions." The Modification Protocol must specify, in advance, the data and methods used to validate future model updates.
If your reference standard and annotation process are written into an authorised PCCP, the sponsor is regulatorily committed to re-running validation against that specification for the life of the device. That converts a one-off dataset sale into a versioned, contracted, recurring evaluation service. The go-to-market instruction that follows is unusually concrete: target the Modification Protocol, not the initial 510(k) dataset.
Now the honest scope limit
This binds devices. The whole apparatus above hangs off the device definition, and Section 520(o)(1)(E) of the FD&C Act, 21 U.S.C. §360j(o)(1)(E), excludes certain clinical decision support software from it (US Code). Two rules matter commercially. Touch a medical image or a signal from an IVD or signal-acquisition system and you are a device regardless of transparency — so radiology, pathology, ECG and continuous monitoring AI are in. The exclusion otherwise requires all three of limbs (i)–(iii), and limb (iii) — that the clinician can "independently review the basis for such recommendations" and is not intended to "rely primarily" on them — is the one most products fail.
But a general-purpose LLM assistant, and the frontier lab training it, sits outside all of this. No regulator is forcing annotation rigour on a chatbot. Those buyers purchase on price, treat data as fungible, and have no deficiency letter waiting for them — and they are where most of the current spend is. That is the tension at the centre of this business: the regulated buyer justifies the expensive pipeline, and the unregulated buyer pays the bills while you build it. The claim that device-destined data commands an order-of-magnitude premium for the same clinician-hours is a reasoned inference from the regulatory structure, not a measured market observation [WEAK]; test it in three discovery conversations before pricing on it.
One further passage is a direct warning to a Europe-based operator. FDA cautions on "the use of data collected outside the U.S. (OUS) in training, which may bias the model if the OUS population does not reflect the U.S. population due to differences in demographics, practice of medicine, or standard of care." A European clinician panel selling into US submissions must answer that structurally: document panel composition, state which standard of care each case is written against, and be able to offer a US-licensed adjudication layer. Handled well it is a differentiator — EU-standard-of-care data for EU submissions, US-adjudicated data for FDA. Handled badly it is a disqualifier. Europe prefers the design you were going to build takes the European half; Clinical reasoning is where the two standards diverge most visibly.