Social media platform X, formally known as Twitter, is facing the heat from UK and EU data regulators after it was found to be potentially in breach of data protection laws for its AI data scraping tactics.
X is using user posts made on its platform to train its upcoming AI model, Grok, but its consent methods are lacking, it has emerged.
Users are automatically opted in to consenting to their data being used to train AI, a practice in breach of UK and EU GDPR.
On X’s settings, Grok can train on users’ posts by default, with a box saying “Allow your posts as well as your interactions, inputs, and results with Grok be used for training and fine” automatically turned on.
While users can turn this feature off via the web platform, GDRP in the UK and EU explicitly says that users must be asked for consent to their data being used for AI training.
Companies are not allowed to use any type of ‘default consent’ under GDPR, including pre-ticked boxes for data sharing.
Data regulators in the UK and EU provided swift statements in relation to the apparent breach, with the Information Commissioner’s Office (ICO) saying that inquiries with X were underway.
“Platforms seeking to use their users’ data to train their AI foundation models must be transparent about their activities,” an ICO spokesperson said.
“They should take steps to proactively notify users well in advance of using data for these purposes, and provide people with ample time and simple processes to object to having their data used in this way.”
Similarly, the Data Protection Commission of Ireland, a subsidiary of the EU’s Data Protection Board, said they were “surprised” by X’s settings, as the data regulator had already been in conversation with the company about its AI and data collection methods.
The Irish DPC spoke with X as soon as the day before the news broke, and have since followed up since the revelation.
Recommended readin
- Demystifying GDPR & AI: Safeguarding Personal Data in the Age of LLMs
- Meta Delays AI Training in Europe Following Regulatory Concern
- OpenAI Changes Data Controller in Bid to Adhere to GDPR
Data scraping, which is often used by AI developers to training their models, involves the scanning of the web and various, mostly public, platforms, for data to train large language models (LLMs).
LLM’s behemoth appetite for data has caused public concern as users become more aware of how large companies are using their personal data to for AI training.
As companies continue to train larger, more advanced models, some, such as Meta, have taken pause on how they access training data for fear of changing regulations as governing bodies grapple with the new normal the age of AI has ushered in.





