Anonymization
Follow topic LLM context A cited markdown file you can paste into your AI assistant (ChatGPT, Claude, a RAG or project knowledge base) to ground it in this topic. Contains: the overview, key law text, case law, enforcement and guidance for this topic. Everything links back to its source on overview.legal — legal information, not advice.Processing anonymized data that cannot be re-identified
Overview
19 sources · Jul 23, 2026Legal Framework
Anonymization sits at the boundary of data protection law: data that is genuinely anonymous falls outside the GDPR entirely, while pseudonymized data remains fully subject to it. The distinction turns on whether a data subject can be re-identified, directly or indirectly, by any reasonably likely means. Article 4(5) GDPR defines pseudonymization as processing personal data so they can no longer be attributed to a specific data subject without the use of additional information, which must be kept separately and subject to technical and organizational measures. True anonymization, by contrast, renders re-identification impossible by any party, using any means reasonably likely to be used — a significantly higher threshold.
The AI Act reinforces these principles. Recital 69 requires providers to implement data minimization and data protection by design and by default throughout the AI system lifecycle, expressly naming anonymization and encryption as necessary measures. Recital 61 extends similar obligations to high-risk AI systems used in judicial and democratic contexts, where bias and opacity risks demand robust safeguards. The Digital Services Act, in Recital 98, similarly treats aggregated, publicly accessible data as a tool for systemic risk research — but only where individual re-identification is effectively precluded.
Key Developments
The CJEU's Planet49 decision established a critical practical benchmark: cookie data linked to a registration number that can be cross-referenced with a user's name and address is personal data, not anonymous data. The mere theoretical possibility of linking identifiers to individuals suffices to bring data within the GDPR's scope. Dutch courts have applied similar reasoning. In the Stichting Benchmark GGZ case before the Rechtbank Midden-Nederland, the court scrutinized whether healthcare benchmark data could be considered sufficiently anonymized, focusing on the re-identification risk inherent in detailed treatment trajectory records.
The Gerechtshof 's-Hertogenbosch (paragraph 4.32) demonstrated the operational side of anonymization, ordering court clerks to produce anonymized copies of judgments and hearing records while preserving the substantive content — illustrating that anonymization must be functional, not merely cosmetic. Similarly, in the covert surveillance case against the Municipality of Delft, the court required black-lining of names, addresses, ages, and phone numbers of NTA employees and respondents, while preserving identifiable findings through labels such as "[NTA employee]" — showing courts expect granular, context-specific anonymization rather than blanket redaction.
Enforcement actions confirm the financial stakes. The Czech DPA fined Avast €13.9 million for disclosing data of approximately 100 million users that the company treated as anonymized but which proved re-identifiable. CNIL imposed €800,000 on Cegedim Santé for transferring customer data without adequate anonymization safeguards.
Practical Guidance
Assess re-identification risk contextually, not abstractly. Planet49 establishes that even indirect linkage through a registration number brings data within the GDPR. Map all reasonably available datasets and cross-referencing possibilities before claiming anonymity.
Separate and protect any key or mapping table. Under Article 4(5), pseudonymized data remains personal data. If a re-identification key exists anywhere in the organization or with a processor, the data is not anonymous.
Apply anonymization by design in AI systems. Recital 69 of the AI Act treats anonymization as a baseline measure, not an optional add-on. Providers must build it into training pipelines, model outputs, and feedback loops from the outset.
Preserve analytical utility through functional anonymization. Courts expect redaction that maintains the substance of findings or conclusions while stripping identifiers — as the Delft court required with labeled placeholders like "[NTA employee]."
Document the anonymization methodology. The Avast enforcement demonstrates that regulators will scrutinize the technical basis for any anonymity claim. Maintain records of the techniques applied, residual risk assessments, and the reasoning supporting the conclusion that re-identification is not reasonably likely.