Site navigation

Why Apple’s AI Problems Could be a Sign of Things to Come

Graham Turner

,

Apple AI
Apple’s suspension of its AI news alert system has sparked debate on the technology’s role in media.

Apple has disabled its controversial AI-powered notification summarisation feature following widespread criticism and complaints.

The feature, part of Apple Intelligence, condensed news alerts from apps into summaries displayed on users’ lock screens.

However, errors in these AI-generated summaries, including a false claim that Luigi Mangione, accused of killing UnitedHealthcare CEO Brian Thompson, had shot himself, raised concerns about misinformation and its impact on the trustworthiness of news outlets. 

If these were slightly clumsy, but still benign mix-ups or half-truths, it could be chalked down to teething problems for a technology that is still very much in its nascent stages.

However, in a climate where we’re becoming ever more drawn to partisan media (mainstream or otherwise) which often prioritises ideology over accuracy – a phenomenon highlighted by a Stanford University study noting that “the public at large tends to put partisanship over truth when consuming news” – the spread of misinformation and disinformation has become a pressing issue.

The World Economic Forum (WEF) has identified this as the top short-term global risk in its 2025 Global Risks Report, underscoring the critical need for responsible storytelling and dissemination. This brings us back to the Apple Intelligence story, where the misuse of trusted media icons further exacerbates the problem, obfuscating the validity of information and undermining public trust or at worse, compelling audiences to accept misinformation as axiomatic.

Source: BBC News

The BBC highlighted the issue in December, lodging a formal complaint and stating that Apple’s AI-generated alerts, branded with news logos, were misleading readers.

Other inaccuracies involving outlets like Sky News and The Washington Post sparked further backlash, with media organisations and industry commentators warning that the technology was not ready for deployment. Apple’s initial response promised a software update to flag AI-generated summaries, but this did little to quell criticism.

To address the problem, Apple has now suspended the feature for news and entertainment apps.

An Apple spokesperson stated the company is “working on improvements and will make them available in a future software update.”

The story has sparked a wave of discourse around AI’s role in media and more broadly, how we gauge veracity in the news we consume. If AI’s to play a central role in this – which it almost certainly is – then, there are (on the side of the technology as it now as a news-telling medium), two huge obstacles to overcome.

AI Hallucinations

AI hallucinations occur when generative AI systems fabricate facts or information because they lack knowledge of a query.

This issue arises from the nature of AI design, which prioritises generating plausible-sounding content rather than verifying its accuracy. 

Generative AI models, trained on vast datasets containing both accurate and inaccurate information, mimic patterns in this data, often reproducing falsehoods or biases.

Essentially, as they largely are now, these models function like advanced autocomplete tools, predicting the next word or phrase based on observed patterns, which can lead to reasonable-sounding but inaccurate outputs.

The inherent challenge is that these systems are not designed to differentiate between truth and falsehood. Even with highly curated training data, the generative nature of these models means they can combine patterns in ways that create new inaccuracies.

To address this, a study conducted last year by researchers from the University of Oxford introduced a statistical model to detect when large language models (LLMs) are likely to hallucinate. This method identifies whether a model is confident in its response or inventing an answer by distinguishing between uncertainty about content and uncertainty about phrasing.

The study’s author highlighted this breakthrough as overcoming previous limitations but emphasised that the method doesn’t address all AI reliability issues, particularly systematic errors.

Dr. Sebastian Farquhar said: “Our method basically estimates probabilities in meaning-space, or ‘semantic probabilities’… The appeal of this approach is that it uses the LLMs themselves to do this conversion.”

“Semantic uncertainty helps with specific reliability problems, but this is only part of the story. If an LLM makes consistent mistakes, this new method won’t catch that. The most dangerous failures of AI come when a system does something bad but is confident and systematic. There is still a lot of work to do.”

New generations of LLMs, such as OpenAI’s o1 model may go some way to addressing this with its chain-of-thought prompting, but that seems like a discussion worth having when it’s properly out in the wild, and the benefits can be comparatively judged.

AI Model Collapse

The development of LLMs that have given rise to the likes of ChatGPT, Gemini, and Apple Intelligence, has got generative AI to a point where it can easily and intuitively create coherent text that is now widely used, for better or worse.

While AI hallucination is a problem we’re well appraised of at this point, there’s another potential looming failure point that’s arguably not been discussed enough.

A study by Nature, released in July of 2024, titled AI models collapse when trained on recursively generated data, highlights a potential fail state that could jeopardise future advancements: “model collapse.”

Model collapse occurs when AI models are trained using data that includes content generated by other AI models. Over time, this leads to the loss of the original data’s diversity and richness, especially the rare or unique details – essentially, it’s a dilution of the original human input.

The study claims that this degradation is irreversible and can happen in various AI models, including LLMs, variational autoencoders (VAEs), and Gaussian mixture models (GMMs), the latter two both being types of machine learning models used for different purposes, primarily generative modeling and clustering.

Currently, LLMs are trained predominantly on human-created text. However, as AI-generated content becomes more prevalent online, future models will inevitably train on this AI-generated content.

This poses a risk: models might start “forgetting” the diversity of human language, making them less accurate and reliable – essentially, the models will become echo chambers of their own homogeny and inaccuracies to the point where it becomes impossible to fix or retrain the models.

Seeking to prove this, the study’s authors conducted experiments to understand the effects of training models on AI-generated data. They found that when models learn from other models’ outputs, their performance degrades significantly. This “model collapse” can lead to a point where models produce only a narrow, less varied range of outputs, losing any unique elements of the original data.

For example, in their experiments with language models, researchers observed that over generations, the models started producing more repetitive and less diverse text. This not only affects the quality of the generated text, but also undermines the model’s ability to understand and generate complex language.


Recommended reading


To prevent this, it’s crucial to maintain access to diverse and original human-generated data. This could involve distinguishing between human and AI-generated content and ensuring that training datasets include a significant proportion of human-created text.

The study also suggests that collaborative efforts among AI developers could help track and manage the provenance of online content.

Without such measures, the ability to train effective and unbiased future AI models could be at risk.

While it’s important to note that this doesn’t have a bearing on the Apple Intelligence situation, as it currently stands, it seems prudent that we stop approaching AI’s shortcomings in a reactionary way and take a view of what could be on the horizon to mitigate it now.

The suspension of Apple’s AI-powered summarisation feature underscores the critical challenges in deploying generative AI responsibly, particularly in sensitive domains like news dissemination.

Issues such as AI hallucinations and the potentially looming threat of model collapse highlight the complexities of ensuring accuracy, reliability, and trustworthiness in AI outputs.

Mix this in with a global audience that’s ever more leaning towards news that reinforces a kind of damaging confirmation bias that colours our views on…well, anything and everything, and it’s hard not to catastrophise and feel like we’re approaching an event horizon for news dissemination.

Graham Turner

Sub Editor

Latest News

AI

Nvidia Launches Open Secure AI Alliance for AI Safety and Security

AI Business Recruitment

Nearly a Quarter of Orgs Reducing Entry-level Hiring Due to AI Automation

Business

Scottish Businesses Turn to Self-funding as Growth Confidence Dips in H2

Data Finance

Payment Leaders are Struggling to Get Real-time Data