Personal data, generated: Generative AI through the lens of Article 4 GDPR – A position paper
Abstract
Personal data is one of the central objects of protection under the General Data Protection Regulation (GDPR). Invoked as early as in the first Recital of the aforementioned normative act, according to which "(...) protection of natural persons in relation to the processing of personal data is a fundamental right", the concept of personal data is one of the cornerstones in the axiological grid of the post-GDPR normative system. Because of that, it should not come as a surprise that the question of what can constitute personal data pursuant to Article 4(1) of the Regulation has been fiercely debated over the last decade, among the legal and non-legal scholars alike. Over those years, a lot has been said about the notion of information, the relationship between that information and a natural person's identifiability and the reasonable measures under which that identification can be performed -- all the necessary building blocks to materialise and interpret an object as personal data under the GDPR. Paradoxically, not a lot has been said about the existence and origin of the object itself, i.e., the situation where information is hypothetically related to an identified or identifiable natural person, but the source of the information is not a factual or fictional claim, but a combination of random characters observed externally - that just by pure coincidence can be linked to an existing natural person, either matching their records fully or partially. Similarly, there is no consensus on what a concatenation of personal data could possibly imply -- and whether such concatenation, that does not fully match a single data subject, could be potentially the object of protection. As we show in this short statement paper, the problem of coincidental generation remains still ambiguous under the current European data protection regime.