Miju Labs

The security dossier

Europe prefers the design you were going to build

GDPR Article 89(1) says purposes that can be met without identifying data subjects 'shall be fulfilled in that manner' — and AI Act Article 10(5), preserved and extended by Regulation (EU) 2026/1744, permits real special-category data only where the job 'cannot be effectively fulfilled by processing other data, including synthetic or anonymised data'.

high confidence9 minupdated 2026-08-30gdpr · ai act · ehds · digital omnibus · europe · synthetic data

A European seat looks like a handicap in a market whose money is currently American. It is not. On the specific architecture this dossier recommends — de-novo cases, blinded adjudication, no patient records — European law is not an obstacle to be tolerated. It is the strongest available third-party endorsement of the design.

Three provisions do the work, and each one says a version of the same thing: if you can do it without real patient data, you must.

GDPR Article 89(1) arguably requires the PHI-free design

Start with the prohibition. Article 9(1) GDPR states that processing of "data concerning health" "shall be prohibited" (Art. 9). This is a prohibition with exhaustive derogations in Article 9(2), not a balancing test — you need both an Article 6 legal basis and an Article 9(2) derogation, and the candidates are worse than founders hope. Explicit consent under 9(2)(a) is withdrawable at any time and does not scale to a corpus. 9(2)(h) covers "preventive or occupational medicine… medical diagnosis, the provision of health or social care" — a care-delivery basis, not a product-development one, and stretching it is a mistake. 9(2)(j), the research derogation, is not self-executing: it requires a Union or Member State law basis plus Article 89(1) safeguards.

And Article 9(4) preserves Member State power to "maintain or introduce further conditions, including limitations" on health data. The regime is not harmonised. A pan-European clinician panel touching real patient records faces up to 27 different answers, each needing local counsel [WEAK] on any individual Member State's specifics. That fragmentation is the practical reason most European health-data plays stall, and it is worth reading alongside Where the supply can legally live, because the same borders that fragment the law also fragment the labour pool.

Now the sentence that matters. Article 89(1) requires safeguards ensuring respect for data minimisation, and closes (Art. 89):

In their words

"Where those purposes can be fulfilled by further processing which does not permit or no longer permits the identification of data subjects, those purposes shall be fulfilled in that manner."

"Shall," not "may." Where the purpose is achievable without identifying data subjects, the non-identifying route is not merely permitted — it is mandated. An operator that can produce its evaluation corpus from de-novo cases and chooses to process real patient records instead is on the wrong side of that sentence. This is the rare case where the compliance-minimal design and the legally preferred design are the same design (Build it without ever touching a patient record).

Two GDPR obligations survive the PHI-free design anyway

The clinicians are data subjects. Their names, licence numbers, specialties, employers, ratings and productivity metrics are personal data requiring a lawful basis, a transparency notice, a retention schedule and access rights. And engaging clinicians as contractors across Member States pulls in employment-adjacent data protection and local worker-representation rules. Budget for a real data-protection function on day one; the saving is in scope, not in existence.

The AI Act makes synthetic data the legally preferred route

Regulation (EU) 2024/1689, Article 6(1), classifies as high-risk an AI system that is, or is a safety component of, a product covered by Annex I Union harmonisation legislation requiring third-party conformity assessment (Art. 6). The MDR (Regulation (EU) 2017/745) and IVDR are Annex I Section A legislation. So any AI-enabled medical device above Class I is a high-risk AI system automatically — MDR classification does the work, and there is no separate risk assessment and no escape hatch.

Article 10 is where that becomes demand (Art. 10). It requires documented governance of "relevant data-preparation processing operations, such as annotation, labelling, cleaning, updating, enrichment and aggregation" (10(2)(c)); examination for biases "likely to affect the health and safety of persons" and measures to detect, prevent and mitigate them (10(2)(f)-(g)); "the identification of relevant data gaps or shortcomings" and how they will be addressed (10(2)(h)); and datasets that are "relevant, sufficiently representative, and to the best extent possible, free of errors and complete" (10(3)).

Annotation practice is named in the statute. Combined with what The regulator wrote your product spec describes, the same evidence package serves an FDA submission and an AI Act technical file — the transatlantic demand story is coherent rather than jurisdiction-specific.

Then Article 10(5), which permits providers of high-risk systems to "exceptionally process special categories of personal data" where strictly necessary for bias detection and correction — but only where, per condition (a):

In their words

"the bias detection and correction cannot be effectively fulfilled by processing other data, including synthetic or anonymised data"

This is a legislative endorsement of the product

Synthetic and anonymised data are the default route; real special-category data is the fallback, available only where the synthetic route demonstrably fails. A vendor supplying credible synthetic bias-probe sets is not selling a convenience — it is selling the thing that removes the buyer's need to process real patient data at all, and therefore the thing that keeps their bias work lawful.

The Digital Omnibus moved the dates, and it is not the setback it looks like

Regulation (EU) 2026/1744 of the European Parliament and of the Council of 8 July 2026, amending Regulation (EU) 2024/1689 as regards simplification of the implementation of harmonised rules on artificial intelligence — the Digital Omnibus on AI — was published in the OJ L series on 24 July 2026 (EUR-Lex, ELI http://data.europa.eu/eli/reg/2026/1744/oj).

Recital 181 gives the rationale — "the delayed availability of standards, common specifications, and alternative guidance and the delayed establishment of national competent authorities" — and the operative amendment to Article 113 sets the new dates:

High-risk routeOriginal AI ActAfter Reg. (EU) 2026/1744
Annex III standalone (Art. 6(2))2 August 20262 December 2027
Annex I embedded — MDR/IVDR devices (Art. 6(1))2 August 20272 August 2028

The naive reading is that demand has been deferred by a year. Three things say otherwise. An uncertain deadline became a budgetable one, and procurement follows certainty rather than urgency — a device sponsor with a firm 2 August 2028 date can put a data line in a 2027 budget, which they could not do while the date was contested. The delay was granted because the standards were not ready, and those standards, when they land, will specify data and annotation practice in exactly the operational detail Article 10 leaves open — a second wave of concrete requirements. And the FDA timeline is untouched, so near-term revenue was never EU-dependent. Plan for US demand through 2026–27 and EU demand ramping into 2028.

The omnibus also inserts a new Article 4a extending the special-category-data basis for bias detection and correction beyond providers of high-risk systems to deployers of high-risk systems and to providers and deployers of other AI systems and models. Critically, Article 4a(1)(a) retains the condition that the work "cannot be effectively fulfilled by processing other data, including synthetic or anonymised data." The synthetic-first preference was not merely preserved; it was extended to a far wider population of actors — including, on its face, deployers who are nowhere near a device submission.

The wider Digital Omnibus is not verified here

Reg. (EU) 2026/1744 is the AI-specific omnibus. The broader package touching GDPR, ePrivacy and the Data Act was a separate proposal whose status could not be confirmed to primary-source standard [UNVERIFIED]. Assume nothing about GDPR amendments until OJ text exists. The same discipline about reading the actual instrument rather than the commentary is the point of The terms of service bite first.

The European Health Data Space is real, and forbids the obvious model

Regulation (EU) 2025/327 (EUR-Lex) creates a statutory route to real European health data for secondary use. Article 53(1)(e) covers scientific research contributing to public health or ensuring quality and safety of healthcare, medicinal products or medical devices, expressly including:

In their words

"(i) development and innovation activities for products or services; (ii) training, testing and evaluation of algorithms, including in medical devices, in vitro diagnostic medical devices, AI systems and digital health applications;"

That is an explicit statutory permission to access real European health data to train and evaluate medical AI — and unlike Article 53(2), which reserves public health, policymaking and official statistics to public sector bodies, point (e) is not reserved. Commercial actors are in scope. The opportunity is genuine.

Then the constraint that decides the business model. Article 73(1) requires access "only through a secure processing environment," and Article 73(2) is categorical: health data access bodies must ensure health data users are "only able to download non-personal electronic health data, including electronic health data in an anonymised statistical format," from that environment.

You can never acquire and resell an EHDS corpus

The data does not come out. What you can do is bring your method, your protocol and your named clinicians inside someone else's environment and carry out anonymised or model-level outputs. EHDS is therefore necessarily a methodology-and-services business, not a data-resale business — which is the shape Build it without ever touching a patient record recommends for entirely unrelated reasons. Every clinician annotator must be named in the data permit; Article 54 confines processing to the permit's stated purposes and the recitals bar making data available to third parties not named in it.

Three more facts price the opportunity. Timing: Chapter IV applies from 26 March 2029, with Member State designations of access bodies due 26 March 2027 and the Commission's secure-processing-environment implementing acts due the same date; some data categories run to 2031 and beyond. This is not a founding-year plan. It is a reason to be European with relationships in place by 2027–28. Forced publication: health data users "shall make public the results or output of secondary use… within 18 months of the completion of the processing" [WEAK] on the article number, verified as Article 71(4) in the text but not re-confirmed — an 18-month disclosure clock is a serious problem for a buyer who wanted a proprietary held-out evaluation set, and must be priced and contracted for explicitly. Fragmentation: Member States may designate one or more bodies, with national ethics review layered on where required, so expect 27 queues and 27 interpretations for some years after 2029 [WEAK].

For Genomics and variant curation in particular, EHDS is the only realistic route to European population-scale data, and the download prohibition makes it a services engagement by construction. Build the panel and the protocol library now; deploy them inside the environments later.