Data governance requirements for high-risk AI training data (Article 10)
In short
Article 10 governs the training, validation, and testing datasets behind a high-risk AI system. It is one of the most technically demanding provisions in the Act, and one of the most consequential, poor data governance is the most common root cause of discriminatory or unreliable AI outputs, and regulators increasingly treat data governance failures as evidence of insufficient diligence in their own right.
Does this apply if the system doesn't train on your data?
Article 10(6) narrows the obligation for systems that use techniques not involving model training, a fixed rule-based system or one that only uses a pre-built model without further training on your own data is subject only to the testing-data requirements, not the full training-data governance set below.
What Article 10(2) requires for datasets that are used
- Documenting relevant design choices, the data’s origin, and (for personal data) the original purpose it was collected for.
- Recording relevant data preparation operations: annotation, labelling, cleaning, updating, enrichment, and aggregation.
- Formulating assumptions about what the data measures and represents.
- Assessing the availability, quantity, and suitability of the datasets needed.
- Examining the data for possible biases likely to affect health, safety, or fundamental rights, or lead to discrimination, and putting measures in place to detect, prevent, and mitigate them.
- Identifying relevant data gaps or shortcomings and how they are addressed.
Data quality, not just data governance process
Beyond the governance process itself, Article 10(3)-(4) requires the datasets to actually be relevant, sufficiently representative, and (to the best extent possible) free of errors and complete, taking into account the specific geographic, contextual, behavioural, or functional setting the system is intended to operate in. A well-documented process around low-quality data does not satisfy the article; both the process and the data quality outcome matter.
The narrow exception for processing special category data
Bias correction sometimes requires processing special category data (race, health status, and similar categories) that would otherwise be restricted under data protection law. Article 10(5) permits this only where strictly necessary, and only with a full safeguard set: confirming no alternative (including synthetic or anonymised data) would work, applying pseudonymisation and strict access controls, never transmitting the data onward, deleting it once the bias is corrected or a retention limit is reached, and documenting the necessity assessment. This is not a general licence to process sensitive data for “fairness” purposes.
Who this falls on
Data governance is a provider obligation, tied to development of the system. A deployer that does not train or fine-tune the system on its own data does not carry this obligation directly, but should still confirm (as part of reviewing the provider’s technical documentation under Article 11) that the underlying dataset practices were addressed, since a deployer inherits the practical consequences of a biased system even without carrying the Article 10 duty itself.
Because “examine for bias” and “sufficiently representative” are outcome standards rather than fixed checklists, the evidence that actually protects an organisation is a documented, repeatable data governance framework, not a one-time bias audit performed before the system launched and never revisited.
Frequently asked questions
- Does Article 10 apply if the system isn't trained on your data?
- Only partly. Article 10(6) narrows the obligation for systems using techniques that don't involve model training, a fixed rule-based system, or one using a pre-built model without further training on your data, is subject only to the testing-data requirements, not the full training-data governance set.
- Is a documented data-governance process enough under Article 10?
- No. Article 10(3)-(4) also requires the datasets themselves to be relevant, sufficiently representative, and (to the best extent possible) free of errors and complete for the system's intended setting. Both the governance process and the data-quality outcome must be satisfied.
- Can you process sensitive personal data to correct bias?
- Only narrowly. Article 10(5) permits processing special-category data for bias correction only where strictly necessary and with full safeguards: confirming no alternative works, applying pseudonymisation and strict access controls, not transmitting the data onward, deleting it once bias is corrected, and documenting the necessity assessment. It is not a general licence to process sensitive data for fairness.
Related guides
Not sure where your company stands?
Our free assessment gives you an indicative result in minutes (free and anonymous) no account needed.
Start the free check