What does AI Act Article 10 (data governance) require?
In short: Article 10 requires high-risk AI training, validation, and test data to be governed: relevant, sufficiently representative, examined for bias, and managed under documented practices covering provenance and preparation. For medical AI, this formalises what good ML development and MDR clinical evidence expectations already push toward.
What the article actually requires
Article 10 sets out data governance and quality requirements for the datasets used to train, validate, and test high-risk AI systems. The data has to be relevant to the intended purpose, sufficiently representative of the population and context the system will actually be used in, examined for possible biases likely to affect health, safety, or fundamental rights, or lead to discrimination, and — to the extent possible — free of errors and complete. Beyond the data itself, you need documented governance practices covering how data was sourced, what design choices and assumptions were made in collecting and labeling it, and what preparation and cleaning steps were applied before training. One provision worth knowing: Article 10(5) permits processing special categories of personal data — health data included — specifically for bias detection and correction, under strict safeguards; see Does GDPR apply to my medical device data? for how that sits with data protection law.
Why this feels familiar if you're already MDR-focused
For teams building medical AI, Article 10 is less a new burden than a formalization of practices already expected under MDR's clinical evidence and risk management requirements — representative training data and documented data provenance are exactly the kind of thing a Notified Body already probes when reviewing a clinical evaluation report for an AI-containing device. The Article 10 documentation slots naturally alongside your existing evidence rather than requiring a parallel data-governance exercise built from nothing — and for AI that is a medical device, it's examined inside your MDR conformity assessment, not in a separate review. As a high-risk obligation, its application date is also subject to the Digital Omnibus deferral — see When do the AI Act deadlines hit medical devices? for the current dates.
Where teams actually get caught out
The most common gap isn't a lack of data quality work — most serious ML teams already do meaningful data curation — it's a lack of documentation proving the work happened in a structured, auditable way. Bias examination that happened informally during model development, with no written record of what was checked and what was found, doesn't satisfy Article 10's documentation expectation even if the underlying data genuinely is representative. Building the documentation habit alongside the data work, not reconstructing it retroactively before a submission, is the difference that matters.
Where next: One Notified Body for MDR and the AI Act · Do I need a separate conformity assessment for the AI Act?
Talk to us about your data governance documentation. Book an expert conversation →
The full guide to data and AI governance for health tech covers this question in context.