Skip to content
Topic Contested in court

Training Data Requirements

LLM context A cited markdown file you can paste into your AI assistant (ChatGPT, Claude, a RAG or project knowledge base) to ground it in this topic. Contains: the overview, key law text, case law, enforcement and guidance for this topic. Everything links back to its source on overview.legal — legal information, not advice.

The AI Act specifically addresses requirements for training, validation, and test data used in high-risk AI systems. This warrants a dedicated topic covering data sourcing, quality standards, documentation requirements, and characteristics that must be maintained for AI model development.

21 linked items 6 Laws4 Guidance1 News10 Literature

Overview

23 sources · Jul 23, 2026

Legal Framework

The AI Act establishes specific requirements for training, validation, and testing data used in high-risk AI systems. Recital 76 identifies training datasets as AI-specific assets vulnerable to cybersecurity threats such as data poisoning, requiring providers to implement protective measures appropriate to the risks. For general-purpose AI model providers, Recital 108 mandates a copyright compliance policy and public disclosure of a summary of content used for training, with the AI Office monitoring compliance without conducting work-by-work copyright assessments. Recital 109 establishes that these obligations must be proportionate to provider size, offering simplified pathways for SMEs and startups while excluding non-professional and scientific research from mandatory compliance.

Where training data contains personal data, the GDPR's data quality principles under Article 6 apply in full. Personal data must be processed fairly and lawfully, collected for specified purposes, adequate and relevant, and kept accurate. The doctrinal analysis confirms that consent must be freely given — data subjects must have a genuine choice and the ability to refuse or withdraw without adverse consequences, with separate consent required for distinct processing operations.

Key Developments

The Court of Justice's ruling in Google Spain established that legitimate interest as a lawful basis requires a balancing test between the controller's interests and the data subject's fundamental rights, with data sensitivity and the passage of time weighing against continued processing. In Peter Puškár, the Court confirmed that the lawful basis criteria are exhaustive — no processing may occur outside these grounds.

The Dutch court decision in the Onderwerp case demonstrates the practical standard for data sourcing transparency: organizations must clearly document how data sources were selected, which sources were searched, how relevance was determined, and how individual documents were assessed. Insufficiently motivated data collection decisions will not survive judicial review.

Practical Guidance

  • Implement a documented copyright compliance policy covering all training data sourcing, and publish a summary of content used for training as required under Recital 108 — the AI Office will monitor compliance without conducting individual work assessments.
  • Apply cybersecurity measures specifically targeting training data vulnerabilities, including protections against data poisoning and adversarial attacks, calibrated to each system's risk profile per Recital 76.
  • Where personal data appears in training sets, ensure each processing operation rests on a valid Article 6(1) GDPR lawful basis, with consent obtained freely and separately for each distinct purpose — bundled consent does not satisfy the requirement.
  • Conduct and document a legitimate interest balancing test for each category of training data, weighing controller objectives against data subject interests, accounting for data sensitivity and temporal factors per the Google Spain standard.
  • Maintain auditable records of data source selection methodology, relevance criteria, and individual document assessment procedures to meet the transparency standard articulated in the Onderwerp ruling.
Everything on this topic, by type links go to the exact provision / paragraph / section
Laws 6
Art. 3(29) ‘training data’ means data used for training an AI system through fitting its learnable parameters; AI Act Art. 3(31) ‘validation data set’ means a separate data set or part of the training data set, either as a fixed or variable split; AI Act Art. 15(5)(cont)(2) The technical solutions to address AI specific vulnerabilities shall include, where appropriate, measures to prevent, detect, respond to, resolve and … AI Act rec 107 Recital 107 — transparency training data summary AI Act Jun 2024 rec 108 Recital 108 — AI Office copyright compliance monitoring AI Act Jun 2024 rec 109 Recital 109 — proportionate compliance for general-purpose AI providers AI Act Jun 2024 rec 111 Recital 111 — systemic risk classification methodology for general-purpose AI models AI Act Jun 2024 rec 76 Recital 76 — AI system cybersecurity protection measures AI Act Jun 2024 rec 67 Recital 67 — high-quality data governance for AI AI Act Jun 2024
Guidance 4
§79 The accents and variations of human speech are vast. While all VVAs are functional once out of the box, their performance can improve by adjusting the… Guidelines 02/2021 on virtual voice assistants §103 I n the event that the user withdraws his or her consent, the data collected from the user can no longer be used for further training of the model. Ne… Guidelines 02/2021 on virtual voice assistants guidelines on the use of facial recognition technology in the area of law enforcement Guidelines 05/2022 on the use of facial recognition technology in the area of law enforcement EDPB May 2023 guidelines on virtual voice assistants Guidelines 02/2021 on virtual voice assistants EDPB Jul 2021 282024 on certain data protection aspects related to Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models EDPB Dec 2024 of the work undertaken by the chatgpt taskforce Report of the work undertaken by the ChatGPT Taskforce EDPB May 2024
News 1
CNIL Artificial intelligence: the action plan of the CNIL CNIL May 2023
Literature 10
i-lex Perspectives for Open Source AI i-lex Jul 2026 SCRIPTed A Journal of Law Technology & Society General-Purpose AI under the EU AI Act: A Conceptual Allocation of Duties across the Value Chain SCRIPTed A Journal of Law Technology & Society Jun 2026 FR American Journal Of Social Sciences And Humanity Research Regulating Algorithm-Based Contracts: How the Eu Artificial Intelligence Act Is Reshaping Risk Allocation in International B2b Transactions American Journal Of Social Sciences And Humanity Research Jun 2026 Law and Economy Italy’s Artificial Intelligence Act and Global AI Governance: The EU Model’s Practice and Prospects Law and Economy Feb 2026 Cambridge Forum on AI Law and Governance Generative AI and data protection Cambridge Forum on AI Law and Governance Jan 2025 International Journal of Computer Applications A Comparative Analysis of the EU AI Act and the Colorado AI Act: Regulatory Approaches to Artificial Intelligence Governance International Journal of Computer Applications Sep 2024 International Journal of Population Data Science ‘Leading by Science’ through Covid-19: the GDPR Automated Decision-Making International Journal of Population Data Science Feb 2021 Unio - EU Law Journal Privacy vs. business convenience: the Mousse judgment and the future of data protection in the EU Unio - EU Law Journal Jun 2025 International Journal of Law and Societal Studies Balancing Security and Privacy: Analyzing the Effectiveness of EU Digital Surveillance Laws in Criminal Proceedings International Journal of Law and Societal Studies Sep 2025 European Economic Letters (EEL) "From Cookies to Context: Adapting Marketing Strategies in a Cookieless Digital Environment" European Economic Letters (EEL) Jun 2025