Training Data Requirements
Follow topic LLM context A cited markdown file you can paste into your AI assistant (ChatGPT, Claude, a RAG or project knowledge base) to ground it in this topic. Contains: the overview, key law text, case law, enforcement and guidance for this topic. Everything links back to its source on overview.legal — legal information, not advice.The AI Act specifically addresses requirements for training, validation, and test data used in high-risk AI systems. This warrants a dedicated topic covering data sourcing, quality standards, documentation requirements, and characteristics that must be maintained for AI model development.
Overview
23 sources · Jul 23, 2026Legal Framework
The AI Act establishes specific requirements for training, validation, and testing data used in high-risk AI systems. Recital 76 identifies training datasets as AI-specific assets vulnerable to cybersecurity threats such as data poisoning, requiring providers to implement protective measures appropriate to the risks. For general-purpose AI model providers, Recital 108 mandates a copyright compliance policy and public disclosure of a summary of content used for training, with the AI Office monitoring compliance without conducting work-by-work copyright assessments. Recital 109 establishes that these obligations must be proportionate to provider size, offering simplified pathways for SMEs and startups while excluding non-professional and scientific research from mandatory compliance.
Where training data contains personal data, the GDPR's data quality principles under Article 6 apply in full. Personal data must be processed fairly and lawfully, collected for specified purposes, adequate and relevant, and kept accurate. The doctrinal analysis confirms that consent must be freely given — data subjects must have a genuine choice and the ability to refuse or withdraw without adverse consequences, with separate consent required for distinct processing operations.
Key Developments
The Court of Justice's ruling in Google Spain established that legitimate interest as a lawful basis requires a balancing test between the controller's interests and the data subject's fundamental rights, with data sensitivity and the passage of time weighing against continued processing. In Peter Puškár, the Court confirmed that the lawful basis criteria are exhaustive — no processing may occur outside these grounds.
The Dutch court decision in the Onderwerp case demonstrates the practical standard for data sourcing transparency: organizations must clearly document how data sources were selected, which sources were searched, how relevance was determined, and how individual documents were assessed. Insufficiently motivated data collection decisions will not survive judicial review.
Practical Guidance
- Implement a documented copyright compliance policy covering all training data sourcing, and publish a summary of content used for training as required under Recital 108 — the AI Office will monitor compliance without conducting individual work assessments.
- Apply cybersecurity measures specifically targeting training data vulnerabilities, including protections against data poisoning and adversarial attacks, calibrated to each system's risk profile per Recital 76.
- Where personal data appears in training sets, ensure each processing operation rests on a valid Article 6(1) GDPR lawful basis, with consent obtained freely and separately for each distinct purpose — bundled consent does not satisfy the requirement.
- Conduct and document a legitimate interest balancing test for each category of training data, weighing controller objectives against data subject interests, accounting for data sensitivity and temporal factors per the Google Spain standard.
- Maintain auditable records of data source selection methodology, relevance criteria, and individual document assessment procedures to meet the transparency standard articulated in the Onderwerp ruling.