The world’s top AI companies are falling disappointingly short when it comes to emerging global safety standards, with many companies failing to meet their own safety guidelines, a new report shows.
The Future of Life Institute‘s 2025 Winter AI Safety Index report shows a clear divide in the safety performance of the top eight AI firms, though none have yet to break an overall grading of B or above.
The researchers assessed the overall safety of some of the world’s top AI companies, assessing their public statements as well as responses to the Institute’s own survey. The scores are based on 35 distinct safety indicators across six categories.
The safety review ranged from internal tools like watermarking to overall transparency policies, questioning the various AI models’ safety as well as the ethical practices of the companies.
Anthropic and OpenAI managed to score a C+, with Google DeepMind scoring a C, based on an average of scores on their risks assessments, safety frameworks, current harms, governance, information sharing, and existential safety.
xAI, Z.ai, Meta, and DeepSeek, all scored a D, with Alibaba Cloud scoring a D-, with all of them falling short in these categories.
The most obvious gaps exist in risk assessments, safety frameworks, and information sharing. Firms scoring low in these often have limited disclosure with weaker evidence of systematic safety processes, as well as uneven adoption of robust evaluation practices.
However, the existential safety emerged as a major gap across all the firms’ safety efforts, the report found.
While Future of Life Institute found that all of the companies were racing towards Artificial General Intelligence (AGI) or superintelligence, but lack explicit plans for controlling these advances, or at least aligning them with current ethical standards. These risky advancements have therefore been left largely unaddressed, even by the more safety aware firms.
While Anthropic, the maker of Claude, and OpenAI, the ChatGPT maker, find themselves at the top of the class, none of the firms seem to be performing well overall.
Recommended reading
- AI Giants ‘Fundamentally Unprepared’ for AGI Safety
- ChatGPT Search Inches Toward Strict EU Oversight as Users Surge
- Beyond DeepSeek: 3 Critical Questions For The Future of AI
Anthropic however has sustained its leadership with a C+ score of 2.67 (the scoring uses a US GPA system, with an A+ equating to 4.3, and an A equal to 4.0), with its high transparency and well developed safety framework. While it has invested in technical safety research and government commitments, it faltered on its shift toward using users interaction for training by default, and its lack of a human uplift trial in its latest risk-assessment.
Those scoring a D and below face major gaps across the board, though Meta’s new safety framework is shown improvement, with xAI formalising its framework, and Z.ai has said an existential-risk plan is in development.
Overall, while firms have made improvements, each company scores well below the emerging standards from regulations such as the EU AI Code of Practice.
The review found that organisations still fail to meet even the most basic safety requirements, such as independent oversight, transparent threat modelling, measurable thresholds, and clearly defined mitigation triggers.





