In the fast-paced realm of artificial intelligence (AI), the line between innovation and infringement gets stepped over time and time again. Compounding that issue is that for the most part, there is no effective way to have an AI model ‘unlearn’ something.
This means that copyrighted materials, information that misrepresents certain topics, or data that breaches user’s privacy could not be taken out of the AI system
Now, researchers and the University of Texas at Austin claim to have pioneered a solution to this problem. Their “machine unlearning” method allows AI systems to selectively forget or block sensitive information, such as violent imagery or copyrighted content, without sacrificing the entirety of their learned knowledge.
“Previously, the only way to remove problematic content was to scrap everything, start anew, manually take out all that data and retrain the model. Our approach offers the opportunity to do this without having to retrain the model from scratch,” said Radu Marculescu, one of the leaders on the project.
The focus of the research, which was published on the arXiv preprint server, lies predominantly in image-to-image models which undergo transformations based on contextual input. The newly developed machine unlearning algorithm provides these models with the capability to remove flagged content promptly, thereby mitigating legal risks and safeguarding user privacy.
Crucially, human oversight remains integral to the process, with dedicated teams tasked with content moderation and removal. This human-AI collaboration ensures an additional layer of scrutiny and responsiveness to user feedback.
“If we want to make generative AI models useful for commercial purposes, this is a step we need to build in, the ability to ensure that we’re not breaking copyright laws or abusing personal information or using harmful content,” said Guihong Li, a graduate research assistant in Marculescu’s lab who worked on the project as an intern.
Recommended reading
- OpenAI Claims AI Training “Impossible” Without Copyrighted Material
- Microsoft to Back Users on AI Copyright Claims
- New Copyright Class-action Lawsuits Filed Against OpenAI and Meta
As AI continues to permeate various facets of society, the development of robust mechanisms to address copyright and privacy concerns is paramount. The work by researchers at The University of Texas at Austin marks a significant step forward in ensuring the responsible and ethical deployment of AI technologies in commercial and societal contexts.
This development comes on the backdrop of a high-profile case where The New York Times sued OpenAI, maker of ChatGPT, saying that the company illegally used its articles as training data to help the chatbot generate content.





