AI has been growing astronomically across almost every industry in some capacity. At the moment, the AI market size is worth about $100 billion (£81.2bn), and expected to grow twenty-fold by 2030 – to £1.5 trillion.
While many parts of the AI industry remain volatile, such as chip availability and market regulation, a new area of concern is now the availability of high-quality data in coming years.
According to a group of researchers, including one from the University of Aberdeen, the stock of high-quality training data will be exhausted as soon as 2026. “Our work suggests that the current trend of ever-growing ML models that rely on enormous datasets might slow down if data efficiency is not drastically improved or new sources of data become available,” say the researchers.
The significance of high quality data in the training of AI algorithms cannot be overstated. For example, ChatGPT underwent training on 570 gigabytes of text data, or about 300 billion words. AI developers need swatches of high-quality content from diverse sources, like books, articles, and scientific papers.
However, the research points out that online data stocks are growing at a slower pace than the datasets required for AI training. While high quality text data could run out by 2026, low quality language data and image data could follow suit between 2030-2050 and 2030-2060 respectively.
The proposed solution is for AI developers to enhance algorithms for them to use existing data more efficiently. This could also lead to reduced computational power which would make the AI landscape greener.
Recommended
- Britain to Crack Down on Unauthorised AI Data Collection
- AI in Scotland: A Present Based on Data, A Future Hinging on Education
- Report: AI Will Be Used for 94% Of Digital Products by 2028
Synthetic data generation is another option, where a data set is artificially generated and mimics properties of real world data. This particular solution is gaining traction already, and as a result, a global forecast predicts the synthetic data generation market will grow by a CAGR of 45.7% to £1.6 billion in 2028.
Another solution put forward is the use of data beyond the free online space. AI developers are negotiating the use of offline repositories with content owners like News Corp, in a potential shift towards paid content deals for training data. This route could also shift the power dynamic currently facing content creators in the face of AI training, where content creators are rarely, if ever, fairly compensated for their data.





