# Training Data Requirements — legal context bundle

> Curated from overview.legal on 2026-08-22. Canonical page: https://overview.legal/topics/training-data-requirements
> Sources are cited per item. Verify against the official texts before relying on them.

The AI Act specifically addresses requirements for training, validation, and test data used in high-risk AI systems. This warrants a dedicated topic covering data sourcing, quality standards, documentation requirements, and characteristics that must be maintained for AI model development.

## Overview

## Legal Framework

The AI Act establishes specific requirements for training, validation, and testing data used in high-risk AI systems. Recital 76 identifies training datasets as AI-specific assets vulnerable to cybersecurity threats such as data poisoning, requiring providers to implement protective measures appropriate to the risks. For general-purpose AI model providers, Recital 108 mandates a copyright compliance policy and public disclosure of a summary of content used for training, with the AI Office monitoring compliance without conducting work-by-work copyright assessments. Recital 109 establishes that these obligations must be proportionate to provider size, offering simplified pathways for SMEs and startups while excluding non-professional and scientific research from mandatory compliance.

Where training data contains personal data, the GDPR's data quality principles under Article 6 apply in full. Personal data must be processed fairly and lawfully, collected for specified purposes, adequate and relevant, and kept accurate. The doctrinal analysis confirms that consent must be freely given — data subjects must have a genuine choice and the ability to refuse or withdraw without adverse consequences, with separate consent required for distinct processing operations.

## Key Developments

The Court of Justice's ruling in *Google Spain* established that legitimate interest as a lawful basis requires a balancing test between the controller's interests and the data subject's fundamental rights, with data sensitivity and the passage of time weighing against continued processing. In *Peter Puškár*, the Court confirmed that the lawful basis criteria are exhaustive — no processing may occur outside these grounds.

The Dutch court decision in the *Onderwerp* case demonstrates the practical standard for data sourcing transparency: organizations must clearly document how data sources were selected, which sources were searched, how relevance was determined, and how individual documents were assessed. Insufficiently motivated data collection decisions will not survive judicial review.

## Practical Guidance

- Implement a documented copyright compliance policy covering all training data sourcing, and publish a summary of content used for training as required under Recital 108 — the AI Office will monitor compliance without conducting individual work assessments.
- Apply cybersecurity measures specifically targeting training data vulnerabilities, including protections against data poisoning and adversarial attacks, calibrated to each system's risk profile per Recital 76.
- Where personal data appears in training sets, ensure each processing operation rests on a valid Article 6(1) GDPR lawful basis, with consent obtained freely and separately for each distinct purpose — bundled consent does not satisfy the requirement.
- Conduct and document a legitimate interest balancing test for each category of training data, weighing controller objectives against data subject interests, accounting for data sensitivity and temporal factors per the *Google Spain* standard.
- Maintain auditable records of data source selection methodology, relevance criteria, and individual document assessment procedures to meet the transparency standard articulated in the *Onderwerp* ruling.

## Legislation (full text of key provisions)

### Recital 107 — transparency training data summary

*Source: AI Act, aiact-rec-107-en, 2024-06-12 — https://overview.legal/posts/93896*

In order to increase transparency on the data that is used in the pre-training and training of general-purpose AI models, including text and data protected by copyright law, it is adequate that providers of such models draw up and make publicly available a sufficiently detailed summary of the content used for training the general-purpose AI model. While taking into due account the need to protect trade secrets and confidential business information, this summary should be generally comprehensive in its scope instead of technically detailed to facilitate parties with legitimate interests, including copyright holders, to exercise and enforce their rights under Union law, for example by listing the main data collections or sets that went into training the model, such as large private or public databases or data archives, and by providing a narrative explanation about other data sources used. It is appropriate for the AI Office to provide a template for the summary, which should be simple, effective, and allow the provider to provide the required summary in narrative form.

### Recital 108 — AI Office copyright compliance monitoring

*Source: AI Act, aiact-rec-108-en, 2024-06-12 — https://overview.legal/posts/93898*

With regard to the obligations imposed on providers of general-purpose AI models to put in place a policy to comply with Union copyright law and make publicly available a summary of the content used for the training, the AI Office should monitor whether the provider has fulfilled those obligations without verifying or proceeding to a work-by-work assessment of the training data in terms of copyright compliance. This Regulation does not affect the enforcement of copyright rules as provided for under Union law.

### Recital 109 — proportionate compliance for general-purpose AI providers

*Source: AI Act, aiact-rec-109-en, 2024-06-12 — https://overview.legal/posts/93900*

Compliance with the obligations applicable to the providers of general-purpose AI models should be commensurate and proportionate to the type of model provider, excluding the need for compliance for persons who develop or use models for non-professional or scientific research purposes, who should nevertheless be encouraged to voluntarily comply with these requirements. Without prejudice to Union copyright law, compliance with those obligations should take due account of the size of the provider and allow simplified ways of compliance for SMEs, including start-ups, that should not represent an excessive cost and not discourage the use of such models. In the case of a modification or fine-tuning of a model, the obligations for providers of general-purpose AI models should be limited to that modification or fine-tuning, for example by complementing the already existing technical documentation with information on the modifications, including new training data sources, as a means to comply with the value chain obligations provided in this Regulation.

### Recital 111 — systemic risk classification methodology for general-purpose AI models

*Source: AI Act, aiact-rec-111-en, 2024-06-12 — https://overview.legal/posts/93904*

It is appropriate to establish a methodology for the classification of general-purpose AI models as general-purpose AI model with systemic risks. Since systemic risks result from particularly high capabilities, a general-purpose AI model should be considered to present systemic risks if it has high-impact capabilities, evaluated on the basis of appropriate technical tools and methodologies, or significant impact on the internal market due to its reach. High-impact capabilities in general-purpose AI models means capabilities that match or exceed the capabilities recorded in the most advanced general-purpose AI models. The full range of capabilities in a model could be better understood after its placing on the market or when deployers interact with the model. According to the state of the art at the time of entry into force of this Regulation, the cumulative amount of computation used for the training of the general-purpose AI model measured in floating point operations is one of the relevant approximations for model capabilities. The cumulative amount of computation used for training includes the computation used across the activities and methods that are intended to enhance the capabilities of the model prior to deployment, such as pre-training, synthetic data generation and fine-tuning. Therefore, an initial threshold of floating point operations should be set, which, if met by a general-purpose AI model, leads to a presumption that the model is a general-purpose AI model with systemic risks. This threshold should be adjusted over time to reflect technological and industrial changes, such as algorithmic improvements or increased hardware efficiency, and should be supplemented with benchmarks and indicators for model capability. To inform this, the AI Office should engage with the scientific community, industry, civil society and other experts. Thresholds, as well as tools and benchmarks for the assessment of high-impact capabilities, should be strong predictors of generality, its capabilities and associated systemic risk of general-purpose AI models, and could take into account the way the model will be placed on the market or the number of users it may affect. To complement this system, there should be a possibility for the Commission to take individual decisions designating a general-purpose AI model as a general-purpose AI model with systemic risk if it is found that such model has capabilities or an impact equivalent to those captured by the set threshold. That decision should be taken on the basis of an overall assessment of the criteria for the designation of a general-purpose AI model with systemic risk set out in an annex to this Regulation, such as quality or size of the training data set, number of business and end users, its input and output modalities, its level of autonomy and scalability, or the tools it has access to. Upon a reasoned request of a provider whose model has been designated as a general-purpose AI model with systemic risk, the Commission should take the request into account and may decide to reassess whether the general-purpose AI model can still be considered to present systemic risks.

### Recital 76 — AI system cybersecurity protection measures

*Source: AI Act, aiact-rec-76-en, 2024-06-12 — https://overview.legal/posts/93834*

Cybersecurity plays a crucial role in ensuring that AI systems are resilient against attempts to alter their use, behaviour, performance or compromise their security properties by malicious third parties exploiting the system’s vulnerabilities. Cyberattacks against AI systems can leverage AI specific assets, such as training data sets (e.g. data poisoning) or trained models (e.g. adversarial attacks or membership inference), or exploit vulnerabilities in the AI system’s digital assets or the underlying ICT infrastructure. To ensure a level of cybersecurity appropriate to the risks, suitable measures, such as security controls, should therefore be taken by the providers of high-risk AI systems, also taking into account as appropriate the underlying ICT infrastructure.

### Recital 67 — high-quality data governance for AI

*Source: AI Act, aiact-rec-67-en, 2024-06-12 — https://overview.legal/posts/93816*

High-quality data and access to high-quality data plays a vital role in providing structure and in ensuring the performance of many AI systems, especially when techniques involving the training of models are used, with a view to ensure that the high-risk AI system performs as intended and safely and it does not become a source of discrimination prohibited by Union law. High-quality data sets for training, validation and testing require the implementation of appropriate data governance and management practices. Data sets for training, validation and testing, including the labels, should be relevant, sufficiently representative, and to the best extent possible free of errors and complete in view of the intended purpose of the system. In order to facilitate compliance with Union data protection law, such as Regulation (EU) 2016/679, data governance and management practices should include, in the case of personal data, transparency about the original purpose of the data collection. The data sets should also have the appropriate statistical properties, including as regards the persons or groups of persons in relation to whom the high-risk AI system is intended to be used, with specific attention to the mitigation of possible biases in the data sets, that are likely to affect the health and safety of persons, have a negative impact on fundamental rights or lead to discrimination prohibited under Union law, especially where data outputs influence inputs for future operations (feedback loops). Biases can for example be inherent in underlying data sets, especially when historical data is being used, or generated when the systems are implemented in real world settings. Results provided by AI systems could be influenced by such inherent biases that are inclined to gradually increase and thereby perpetuate and amplify existing discrimination, in particular for persons belonging to certain vulnerable groups, including racial or ethnic groups. The requirement for the data sets to be to the best extent possible complete and free of errors should not affect the use of privacy-preserving techniques in the context of the development and testing of AI systems. In particular, data sets should take into account, to the extent required by their intended purpose, the features, characteristics or elements that are particular to the specific geographical, contextual, behavioural or functional setting which the AI system is intended to be used. The requirements related to data governance can be complied with by having recourse to third parties that offer certified compliance services including verification of data governance, data set integrity, and data training, validation and testing practices, as far as compliance with the data requirements of this Regulation are ensured.

## Guidance

### Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models

*Source: EDPB, opinion-282024-on-certain-data-protection-aspects-related-to-en, 2024-12-18 — https://overview.legal/posts/125697 — original: https://www.edpb.europa.eu/documents/opinion-of-the-board-art-64/opinion-282024-on-certain-data-protection-aspects-related-to_en*

Adopted 1 Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models Adopted on 17 December 2024 Adopted 2 Executive summary AI technologies create many opportunities and benefits across a wide range of sectors and social activities. By protecting the fundamental right to data protection, GDPR supports these opportunities and promotes other EU fundamental rights, including the right to freedom of thought, expression and information,…

### Report of the work undertaken by the ChatGPT Taskforce

*Source: EDPB, report-of-the-work-undertaken-by-the-chatgpt-taskforce-en, 2024-05-24 — https://overview.legal/posts/125752 — original: https://www.edpb.europa.eu/documents/task-force-report/report-of-the-work-undertaken-by-the-chatgpt-taskforce_en*

Report of the work undertaken by the ChatGPT Taskforce 23 May 2024 Final 2 Final 3 D ISCLAIMER The positions presented in this document result from the coordination of the members of the ChatGPT taskforce with a view to handling investigations regarding the service ChatGPT provided by the US based company OpenAI OpCo, LLC . They reflect the common denominator agreed by the S upervisory A uthorities in their interpretation of the applicable provisions of the GDPR in relation to the matters that…

### Guidelines 05/2022 on the use of facial recognition technology in the area of law enforcement

*Source: EDPB, edpb-guidelines-on-the-use-of-facial-recognition technology-in-the-area-of-law-enforcement, 2023-05-17 — https://overview.legal/posts/38075 — original: https://www.edpb.europa.eu/documents/guideline/guidelines-052022-on-the-use-of-facial-recognition-technology-in-the-area-of_en*

More  and  more  law  enforcement  authorities  (LEAs)  apply  or  intend  to  apply  facial  recognition technology (FRT). It may be used to authenticate or to identify a person and can be applied on videos (e.g. CCTV) or  photographs. It may be used for various purposes, including to search for persons  in police watch lists or to monitor a person's movements in the public space. FRT is  built on the processing of biometric data , therefore, it encompasses the processing of special categories ...

### Guidelines 02/2021 on virtual voice assistants

*Source: EDPB, edpb-guidelines-on-virtual-voice-assistants, 2021-07-07 — https://overview.legal/posts/38077 — original: https://www.edpb.europa.eu/documents/guideline/guidelines-022021-on-virtual-voice-assistants_en*

A virtual voice assistant (VVA) is a service that understands voice commands and executes them or mediates with other IT systems if needed. VVAs are currently available on most smartphones and tablets, traditional computers, and, in the latest years, even standalone devices like smart speakers. VVAs act as interface between users and their computing devices and online services such as search engines  or  online  shops.  Due  to  their  role,  VVAs  have  access  to  a  huge  amount  of  personal...

## Recent developments

### Artificial intelligence: the action plan of the CNIL

*Source: CNIL, 2023-05-16 — https://overview.legal/posts/6206 — original: https://www.cnil.fr/en/artificial-intelligence-action-plan-cnil#entry-5218*

The main thing is:

The CNIL has been undertaking work for several years to anticipate and respond to the issues raised by AI.
In 2023, it will extend its action on augmented cameras and wishes to expand its work to generative AIs, large language models and derived applications (especially chatbots).
Its action plan is structured around four strands:

to understand the functioning of AI systems and their impact on people;
enabling and guiding the development of privacy-friendly AI;
federate and

## Literature

### Perspectives for Open Source AI

*Source: i-lex, 2026-07-07 — https://overview.legal/posts/83512 — original: https://doi.org/10.60923/issn.1825-1927/23382*

The world’s first most comprehensive law regulating artificial intelligence, the EU Artificial Intelligence Act, has been enacted in June 2024 and entered into force in August 2024. The AI Act aims to provide transparency and ensure safe use of AI systems by introducing obligations and requirements for developers and deployers based on the risk posed by AI systems. Despite the long legislation process that launched in 2020 and multiple negotiations, the final version of the Act includes a number

### General-Purpose AI under the EU AI Act: A Conceptual Allocation of Duties across the Value Chain

*Source: SCRIPTed A Journal of Law Technology & Society, 2026-06-30 — https://overview.legal/posts/132370 — original: https://doi.org/10.2218/scrip.12300*

This article examines how the final version of the EU Artificial Intelligence Act (“AI Act”, adopted 2024) allocates obligations across the AI value chain, with a focus on general-purpose AI (“GPAI”) or foundation models. It proposes a taxonomy of key actors – foundation model providers, fine-tuners, integrators, and deployers – and analyses the interfaces between them, including documentation tools (model cards, system cards) and logging requirements. Building on principles of control, foreseea

### Regulating Algorithm-Based Contracts: How the Eu Artificial Intelligence Act Is Reshaping Risk Allocation in International B2b Transactions

*Source: American Journal Of Social Sciences And Humanity Research, 2026-06-22 — https://overview.legal/posts/132566 — original: https://doi.org/10.37547/ajsshr/volume06issue06-20*

This article examines how the EU Artificial Intelligence Act (the “EU AI Act”) transforms the allocation of risk in cross-border B2B contracts built on algorithmic and automated decision-making systems. Drawing on doctrinal and comparative legal analysis of the EU AI Act, related EU instruments - the Model Contractual Clauses for AI Procurement and initiatives of the European Law Institute - and UNCITRAL’s Model Law on Automated Contracting, the study shows that before the EU AI Act, risk distri

### Italy’s Artificial Intelligence Act and Global AI Governance: The EU Model’s Practice and Prospects

*Source: Law and Economy, 2026-02-25 — https://overview.legal/posts/132619 — original: https://doi.org/10.63593/le.2788-7049.2026.03.004*

The Italian Artificial Intelligence Act, enacted on September 17, 2025, represents the first comprehensive national implementation of the European Union’s AI Act. This study examines the Italian legislation through the theoretical lens of multi-level governance, analyzing its dual function as both a “bridging legislation” that translates EU framework into domestic practice and a site of significant regulatory innovation. Through detailed textual analysis and case studies, particularly in healthc

### Generative AI and data protection

*Source: Cambridge Forum on AI Law and Governance, 2025-01-01 — https://overview.legal/posts/83528 — original: https://doi.org/10.1017/cfl.2024.2*

Abstract Generative artificial intelligence (AI) has catapulted into the legal debate through the popular applications ChatGPT, Bard, Dall-E and others. While the predominant focus has hitherto centred on issues of copyright infringement and regulatory strategies, particularly in the context of the AI Act, a critical but often overlooked issue lies in the friction between generative AI and data protection laws. The rise of these technologies highlights unresolved tension between safeguarding fun

## Related topics

- **Data Governance for AI** — https://overview.legal/topics/data-governance-ai
  The AI Act's section on 'Data and data governance' requires specific provisions for managing training data, validation data, and test data in AI systems. This c
- **Artificial Intelligence** — https://overview.legal/topics/ai
  AI systems and their implications for data protection
- **Personal Data** — https://overview.legal/topics/persoonsgegevens
  Information relating to identified or identifiable natural persons
- **AI Value Chain Actors and Roles** — https://overview.legal/topics/ai-value-chain-actors
  The content focuses on responsibilities distributed across different actors in the AI value chain. A dedicated topic for understanding the various actors, their
- **Transparency** — https://overview.legal/topics/transparantie
  Openness about data processing activities
- **Accountability** — https://overview.legal/topics/accountability
  Principle of demonstrating GDPR compliance

---
Generated by overview.legal · https://overview.legal/topics/training-data-requirements · 2026-08-22
