A recent study out of the University of Exeter, published in the journal Big Data and Society, underscores the importance of transparency, accountability, and fairness in handling synthetic data, which is becoming increasingly prevalent in various fields.
Synthetic data is created through machine learning algorithms from real-world data, and offers privacy preserving alternatives to traditional data sources. This is a proven and valuable method in cases where sharing actual data poses sensitivity concerns, scarcity issues, or quality constraints.
However, the study shows that existing data protection laws fall short of regulating the processing of synthetic data comprehensively. Laws like the EU GDPR primarily address personal data, defined as any information relating to an identified or identifiable natural person.
While fully synthetic datasets are exempt from GDPR regulations, exceptions could arise if there is any personal information that could pose a risk of re-identification. This ambiguity surrounding the threshold for re-identification risk complicates the legal landscape and operational aspects of processing synthetic data, according to the report.
“Clear guidelines for all types of synthetic data should be established. They should prioritise transparency, accountability and fairness,” said professor Ana Beduschi, author of the study.
Recommended reading
- Scotland’s Synthetic Data Journey
- ICO Joins Global Data Protection Enforcement Programme
- ICO Publishes New Data Protection Fining Guidance
“Having such guidelines is especially important as generative AI and advanced language models such as DALL-E 3 and GPT-4—which can both be trained on and generate synthetic data—may facilitate the dissemination of misleading information and have detrimental effects on society. Adhering to these principles could thus help mitigate potential harm and encourage responsible innovation,” she continued.
“Accordingly, synthetic data should be clearly labelled as such and that information about its generation should be provided to users.”
By prioritising transparency, accountability, and fairness, these guidelines aim to ensure the responsible use of synthetic data in various domains.





