Training Data Requirements
Follow topic LLM context A cited markdown file you can paste into your AI assistant (ChatGPT, Claude, a RAG or project knowledge base) to ground it in this topic. Contains: the overview, key law text, case law, enforcement and guidance for this topic. Everything links back to its source on overview.legal — legal information, not advice.The AI Act specifically addresses requirements for training, validation, and test data used in high-risk AI systems. This warrants a dedicated topic covering data sourcing, quality standards, documentation requirements, and characteristics that must be maintained for AI model development.
Overview
15 sources · Sep 8, 2026Legal Framework
Training data requirements for AI systems are governed primarily by the AI Act, with overlapping obligations under the GDPR where personal data is involved. Article 15 of the AI Act imposes accuracy and robustness obligations on high-risk AI systems that directly implicate training, validation, and test data quality. Providers must ensure that systems perform consistently throughout their lifecycle, and that feedback loops in continuously learning systems do not propagate bias.
For general-purpose AI models, Recital 107 establishes a transparency obligation: providers must publish a sufficiently detailed summary of the content used for training. This requirement is designed to enable rights holders—including copyright holders—to enforce their interests under Union law.
"it is adequate that providers of such models draw up and make publicly available a sufficiently detailed summary of the content used for training the general-purpose AI model"
— AI Act Recital 107
Where training data includes personal data, the GDPR applies in full. Consent, where relied upon as a lawful basis, must be freely given, specific, and informed. The EDPB has confirmed that withdrawal of consent has immediate consequences for ongoing training:
This distinction—between prospective use of data and retroactive deletion of trained models—creates a practical tension that providers must navigate.
Key Developments
The EDPB has addressed training data quality through its guidance on bias evaluation, highlighting a critical risk in test dataset selection. The concern is not merely data quantity but representativeness:
This guidance signals that regulators will scrutinize whether test datasets carry historical, representation, or measurement bias. A model trained on geographically or temporally narrow data may achieve high benchmark scores while failing in deployment—a discrepancy that regulators view as a data governance deficiency, not merely a performance issue.
The EDPB's voice assistant guidelines further establish that consent withdrawal blocks further training use of an individual's data, though the trained model itself need not be deleted. However, the Board flags that membership inference and reconstruction attacks may allow extraction of personal data from trained models, requiring mitigation measures beyond simple data segregation.
Status of the Debate
This topic is actively contested. The boundary between GDPR lawful processing requirements and AI Act data governance obligations remains unsettled—particularly regarding whether GDPR consent withdrawal retroactively undermines the lawfulness of a model already trained on that data. No court has definitively ruled on whether the "model need not be deleted" principle survives a challenge grounded in Article 17 GDPR (right to erasure) or Article 5(1)(b) (purpose limitation). The interaction between AI Act Recital 107's training data summary requirement and trade secret protection under Directive (EU) 2016/943 also lacks authoritative interpretation. Resolution will likely require a CJEU preliminary reference or coordinated guidance from the AI Office and EDPB.
Practical Guidance
Document data provenance and representativeness: Under Article 15, providers must demonstrate that training, validation, and test datasets are sufficiently representative to achieve appropriate accuracy and robustness throughout the system lifecycle. Maintain records linking data sources to performance metrics.
Implement consent withdrawal mechanisms for training data: Where consent is the lawful basis, ensure technical and organizational processes can isolate and exclude withdrawn data from future training cycles, consistent with EDPB guidance.
Publish a training data summary for GPAI models: Prepare a narrative summary of main data collections and sources used, balancing transparency under Recital 107 with trade secret protection, using the AI Office template once available.
Test for dataset bias and overfitting: Validate that test datasets do not carry historical or representation bias that inflates benchmark performance. Document mitigation measures where geographic, temporal, or demographic gaps exist.
Address feedback loop risks in continuously learning systems: For high-risk AI systems that learn post-deployment, implement technical measures to detect and mitigate biased outputs re-entering the training pipeline, as required by Article 15(4).
Nothing of this type on this topic.