Site navigation

OpenAI Claims AI Training “Impossible” Without Copyrighted Material

Elizabeth Greenberg

,

openAI copyright
The admittance comes after OpenAI is facing two new lawsuits for using copyrighted material to train their AI models. 

In a letter to the House of Lords committee on communications and digital select, OpenAI has admitted that it would be impossible to train their large language models without copyrighted work in order to meet the “needs” of people today.

The admittance was provided in a document of written evidence for the committee’s inquiry into large language model, which asked the creators of GPT-4 and ChatGPT about the future of LLMs, the risk of frontier AI, and regulatory intervention  on the AI landscape.

At the same time, OpenAI has stated that they respect and have complied with all requirements of copyright laws, though they contend that “we believe that legally copyright law does not forbid training.”

The prolific AI company went on to say that they have introduced tools for creators to stop the company from using or accessing their work, exclude their content from their web crawler – a tool that the company uses to scan the internet for data to train their AI – and has created an opt-out process for creators to exclude their work from DALL∙E training material.

OpenAI’s logic for using copyrighted material appears relatively sound, if one accepts the modern nature of the availability of intellectual material – “copyright today covers virtually every sort of human expression,” the letter argued, including photographs, software code, blogs, and government documents.

Limiting AI training models to opensource material and data would therefore not reflect the modern world, nor its diversity in language, experience, and expression, disabling AI systems from meeting “the needs of today’s citizens.”

These “needs” however, appear relative.

And, while OpenAI has introduced certain features enabling creators to block their material from being used as training data, the company has already faced numerous lawsuits from authors, artists, and researchers for using their copyrighted material without prior consent, or offering any compensation.

Further, OpenAI’s opt-out model challenges more progressive data protection regulations, including GDPR, which typically require an opt-in option for harvesting user data that is not integral to the functioning of a service.

At the very least, consumers would need to be made aware that their use of, posts on, and interaction with a website may be used to train an AI model.

For some, OpenAI’s opt-out policy is also too little, too late.

At the start of this year, the New York Times issued a lawsuit against OpenAI and Microsoft for using the work of its journalists to train their LLMs. This is after the media company took advantage of OpenAI’s new policy and blocked their web scraper from using their published material.


Recommended reading


Two more writers also filed a lawsuit just this week in New York alleging the same thing of their nonfiction work, following authors and celebrities like Sarah Silverman and George R.R. Martin.

Further, a class action lawsuit in California claimed the very nature of data scraping to train AIs is invasive, though this case was dismissed later without prejudice, meaning it can be refiled by the plaintiffs.

Despite these continued challenges, Microsoft has been steadfast, even offering to legally back users over any AI copyright claims for using the company’s AI co-pilot feature.

In the letter, OpenAI has put its cards on the table, in a way, and challenged the House of Lords to accept their admittance as fact. Either allow copyrighted material to be used, or say goodbye to advanced LLMs as we have grown to know them.

Whether the government, the public, and creators accept this, however, will remain to be seen.

Elizabeth Greenberg

Staff Writer

Latest News

AI

Nvidia Launches Open Secure AI Alliance for AI Safety and Security

AI Business Recruitment

Nearly a Quarter of Orgs Reducing Entry-level Hiring Due to AI Automation

Business

Scottish Businesses Turn to Self-funding as Growth Confidence Dips in H2

Data Finance

Payment Leaders are Struggling to Get Real-time Data