Digital collections of arts and archives are being threatened by the torrent of bots AI companies have deployed to harvest data, new research shows.
An report from GLAM-E Lab, which studies issues surrounding Galleries, Libraries, Archives, and Museums (GLAMs), shows that these digital spaces are being swarmed by AI bots trawling their websites for data to train their AI models.
The proliferation of bots have gone so far that GLAM institutes claim AI companies do not understand the processing burden these bots have on their sites.
The surges of AI bot traffic can go so far that it kicks collections offline, according to GLAM-E Lab’s research.
An anonymous survey of 43 GLAM organisations revealed heightened alarm concerning the ‘aggressive’ harvesting of content by AI bots across digital collections.
Of the 43 organisations, 39 reported an increase in traffic, with 27 of these attributing this increase to AI training data bots.
“Respondents worry that swarms of AI training data bots will create an environment of unsustainably escalating costs for providing online access to collections,” the report said.
While some of the bots identify themselves and other do not, and while some trace the increase in bots back to 2021 while others only note the increase in bot traffic this year, respondents say that the robots.txt directives they publish to offer guidelines for web publishers are ineffective in stopping swarms of AI bots.
Other bot defence systems, like those offered by AWS and Cloudflare, however, are more effective, though GLAM-E Lab says the the problem is more complex.
Organisations that want their material available to the public may not want to hide materials behind logins, and some bot traffic may be useful, such as indexing bots.
Recommended reading
- UK Data Use and Access Bill Set to Become Law
- How Can Creative Industries Better Navigate AI’s Copyright Challenges?
- AI at Work Sparks Copyright Concerns for UK Firms
The GLAM-E Lab report echoes other investigations and research made by the Confederation of Open Access Repositories, which found that bot traffic is leading to service outages or at least slowdowns of sites.
The Wikimedia Foundation and SourceHut issued similar complaints about the aggressive tactics of these bots in harvesting data to train AI, which is their assumed purpose.
GLAM-E Labs urged AI companies to consider the negative ramifications their swarms of bots can have on the platforms they target.
The research opens up a new avenue of concern as many GLAM institutions battle AI data harvesting over copyright concerns. Getty Images issued a lawsuit directed at Stability AI over their use of copyrighted material.
The UK government secretary of state for science, innovation, and technology Peter Kyle said that the government is working on addressing how UK copyright laws apply to AI, though this has left many questions unanswered. Copyright became a battleground for the new Data Use and Access Bill that was just passed through Parliament, with the UK government saying it will publish a report on AI and copyright instead of hammering the matter out in the new legislation.





