The development of LLMs that have given rise to the likes of GPT-4 has got generative AI to a point where it can easily and intuitively create coherent text that is now widely used, for better or worse.
However, a new study by Nature, titled AI models collapse when trained on recursively generated data highlights a potential problem that could jeopardise future advancements: “model collapse.”
Model collapse occurs when AI models are trained using data that includes content generated by other AI models. Over time, this leads to the loss of the original data’s diversity and richness, especially the rare or unique details – essentially, it’s a dilution of the original human input.
The study claims that this degradation is irreversible and can happen in various AI models, including LLMs, variational autoencoders (VAEs), and Gaussian mixture models (GMMs), the latter two both being types of machine learning models used for different purposes, primarily generative modeling and clustering.
Currently, LLMs are trained predominantly on human-created text. However, as AI-generated content becomes more prevalent online, future models will inevitably train on this AI-generated content.
This poses a risk: models might start “forgetting” the diversity of human language, making them less accurate and reliable – essentially, the models will become echo chambers of their own homogeny and inaccuracies to the point where it becomes impossible to fix or retrain the models.
Seeking to prove this, the study’s authors conducted experiments to understand the effects of training models on AI-generated data. They found that when models learn from other models’ outputs, their performance degrades significantly. This “model collapse” can lead to a point where models produce only a narrow, less varied range of outputs, losing any unique elements of the original data.
For example, in their experiments with language models, researchers observed that over generations, the models started producing more repetitive and less diverse text. This not only affects the quality of the generated text, but also undermines the model’s ability to understand and generate complex language.
Recommended reading
- 75% of Enterprises Push Ahead With AI, Despite Data Governance Concerns
- 64% of People Don’t Want Companies to Use AI in Customer Service
- Comment | What AI Adopters in Scotland Can Learn From the dot-com Bubble
The potential for model collapse raises important questions about the future of AI training.
To prevent this, it’s crucial to maintain access to diverse and original human-generated data. This could involve distinguishing between human and AI-generated content and ensuring that training datasets include a significant proportion of human-created text.
The study also suggests that collaborative efforts among AI developers could help track and manage the provenance of online content.
Without such measures, the ability to train effective and unbiased future AI models could be at risk.





