The two lawsuits, filed by the San Francisco-based Joseph Saveri Law Firm alongside lawyer Matthew Butterick, claim that the copyrighted works of Sarah Silverman, Christopher Golden, and Richard Kadrey were used to train OpenAI’s and Meta’s large language models (LLMs) without the authors’ permission.
As its name implies, an LLM is a type of artificial intelligence (AI) algorithm that utilises massively large datasets and deep learning techniques to then generate a wide variety of text-based outputs, from writing Python code and poems, to summarising a highly complex scientific term or published book.
Large language models are also what underpin contemporary generative AI tools such as OpenAI’s ChatGPT chatbot.
In the new filing against OpenAI, it claims that the authors “did not consent to the use of their copyrighted books as training material for ChatGPT. Nonetheless, their copyrighted materials were ingested and used to train ChatGPT.”
Similarly, the filing challenging Meta — whose LLM is called LLaMA — alledges that the authors “did not consent to the use of their copyrighted books as training material for LLaMA,” and that “Nonetheless, their copyrighted materials were copied and ingested as part of training LLaMA.”
In particular, the latter filing proclaims that “Many of Plaintiffs’ books appear in the Books3 dataset” — a dataset that is supposedly a part of ThePile, an alledgedly 825 gibibyte dataset that contains many smaller datasets combined together.
The former filing also claims that when prompted, ChatGPT “generates summaries of Plaintiffs’ copyrighted works—something only possible if ChatGPT was trained on Plaintiffs’ copyrighted works.”
Recommended
- How Has One of Scotland’s Most Influential Tech Investors Evolved?
- New Report Finds UK Fintech Funding Down 37% in H1 2023
- Microsoft and AWS Push Back Against Cloud Market Investigation
Speaking on the Meta filing in particular, Joseph Saveri, founder of the Joseph Saveri Law Firm, said: “As artificial intelligence continues to change every aspect of the modern world, we must recognize and protect the rights of artists such as these authors against unlawful theft and fraud.
“LLaMA is not just an infringement of authors’ rights; whether they aim to or not, these products will eliminate ‘author’ as a viable career path. This case represents a larger fight for preserving ownership rights for all artists and other creators.”
Last month, the law firm filed another class-action lawsuit against OpenAI regarding copyright for the authors Paul Tremblay and Mona Awad.
The filed lawsuits, no matter what procedes, highlight the broader concerns that many people — not least professionals in the creative industries, as well as global regulators — currently have concerning potential copyright issues amid the generative AI explosion.
According to Joseph Saveri and Matthew Butterick, “Since the release of OpenAI’s ChatGPT system in March 2023, we’ve been hearing from writers, authors, and publishers who are concerned about its uncanny ability to generate text similar to that found in copyrighted textual materials, including thousands of books,” they said.





