Contributing to the EDPB Consultations on Anonymisation and Web Scraping for Generative AI

The rapid development of generative AI continues to challenge established data protection concepts and practices. To contribute to the ongoing European discussion, I recently submitted feedback to the European Data Protection Board (EDPB) on its draft Guidelines 02/2026 on Anonymisation and Guidelines 03/2026 on Web Scraping in the Context of Generative AI. My comments draw on practical experience advising technology companies, research organisations and Horizon Europe projects on data governance, AI governance and privacy compliance.

Anonymisation: supporting innovation through proportionate implementation

My response to the anonymisation consultation welcomes the EDPB's efforts to provide greater clarity on a topic that remains one of the most difficult areas of data protection law. In particular, I encouraged the EDPB to include additional guidance on AI-related use cases, such as foundation models, embeddings, vector databases and synthetic data generation. I also suggested further clarification on how anonymisation principles should apply to AI models and derived representations, where organisations often struggle to assess whether personal data remains present in a meaningful sense. Finally, drawing on practical experience from collaborative research and innovation projects, I advocated for a clearer recognition of proportionality and risk-based accountability. Organisations need guidance not only on whether anonymisation has been achieved, but also on what constitutes a reasonable and proportionate level of assessment in low-risk research, educational and public-engagement activities.

Web scraping and generative AI: improving legal certainty in practice

My response to the web scraping consultation focuses on areas where organisations frequently face legal uncertainty when developing or deploying AI systems. In particular, I encouraged the EDPB to provide a more structured framework for assessing individuals' "reasonable expectations" and to recognise that different categories of publicly available information, such as academic publications, professional biographies, public registers and personal social-media content, may justify different assessments. I also proposed additional guidance for organisations using third-party datasets and AI models, many of whom were not involved in the original scraping activities but nevertheless need to understand their due diligence obligations. Finally, I highlighted the continuing challenges around Article 14 transparency obligations at scale and encouraged further exploration of practical technical mechanisms that could support transparency and accountability in the AI ecosystem.

The full consultations remain open for feedback until 30 October 2026 via the EDPB's public consultation process. The final guidelines are likely to play an important role in shaping how organisations approach anonymisation, web scraping and AI governance across Europe in the coming years.