Research from the Massachusetts Institute of Technology (MIT) demonstrated that machine learning (ML) models trained exclusively with synthetic images outperform their counterparts trained with real images on a large scale.
For the study, synthetic images are created with text-to-image models like Stable Diffusion, and by using a strategy called “multi-positive contrastive learning,” the research team could train ML models using these synthetic, AI-generated images.
“We’re teaching the model to learn high-level concepts through context and variance, not just feeding it data,” said Lijie Fan, MIT Ph.D. student in electrical engineering and lead researcher on the project, speaking with MIT News.
“StableRep goes beyond traditional methods by considering multiple synthetic images as positive pairs, allowing the model to delve deeper into the underlying concepts rather than just pixel-level details.”
This development by MIT extends beyond improved performance. The looming existential threat for AI models of all kinds is that data stock is growing at a slower rate than is required for training AI models.
Recent research shows that stock for text data could be exhausted as soon as 2026, and image data by 2030. The proposed solutions are either to use the existing data more efficiently, or to leverage synthetic data generation.
“The capacity to produce high-calibre, diverse synthetic images on command could help curtail cumbersome expenses and resources associated with traditional data collection methods,” continued Fan.
The team’s enhanced variant, StableRep+, outperformed traditional models not only in accuracy but also in efficiency, using 20 million synthetic images compared to 50 million real images.
Recommended reading
- Scotland’s Synthetic Data Journey
- Microsoft Creating Tool to Weed Out AI Bias
- Smart Data Foundry Provides Synthetic Data to Reduce Fraud
However, the researchers acknowledge challenges, including the slow pace of image generation, semantic mismatches between text prompts and images, potential biases, and complexities in image attribution.
“One dream of generative model learning has long been to be able to generate data useful for discriminative model training,” commented David Fleet, Google DeepMind researcher and University of Toronto professor of computer science, speaking with MIT about the paper.
“While we have seen some signs of life, the dream has been elusive, especially on large-scale complex domains like high-resolution images. This paper provides compelling evidence, for the first time to my knowledge, that the dream is becoming a reality. They show that contrastive learning from massive amounts of synthetic image data can produce representations that outperform those learned from real data at scale, with the potential to improve myriad downstream vision tasks.”





