Yoshua Bengio, considered one of the “godfathers” of modern AI and a pioneer in deep learning, has joined a project funded by the UK Government’s Advanced Research and Invention Agency which aims to construct an AI ‘gatekeeper’, a system tasked with understanding and reducing the risks of other AI agents.
The project, dubbed Safeguarded AI, will receive £59 million over the next four years, having been launched in January 2023 to invest in potentially transformational scientific research.
Bengio, who was awarded the 2018 Turing Award for his work in artificial intelligence, will work with the project on combining scientific world models, essentially simulations of the world, with mathematical proofs. These proofs would include explanations of the AI’s work, and humans will then be tasked with verifying whether the AI model’s safety checks are correct.
Eventually, the AI systems that he and the team build will offer quantitative guarantees, such as risk scores, about the potential effect that other AI models could have on the real world.
Bengio has claimed that, at present, technology companies cannot provide mathematical guarantees that AI systems will behave as intended, which could result in adverse outcomes for those developing and using the technology.
Speaking on the complexity of the issue, and the importance of using AI constructs to police other AI agents, Bengio told MIT Technology Review: “[It’s] the only way, because at some point these AIs are just too complicated. Even the ones that we have now, we can’t really break down their answers into human, understandable sequences of reasoning steps.”
Recommended reading
- UK and South Korea Prepare for Next Global AI Safety Summit
- UK and US Strike New AI Safety Partnership
- UK AI Summit: What It Is, Who’s Going, and What the Tech Industry Wants
Bengio is already an advocate for the importance of establishing a global AI safety framework. He currently chairs the International Scientific Report on the Safety of Advanced AI, as well as being a member of the UN’s Scientific Advisory Board for Independent Advice on Breakthroughs in Science and Technology.
Regarding the appointment of Bengio the official ARIA blog posted: “We’re excited to welcome Professor Yoshua Bengio as Scientific Director for Safeguarded AI.
“Yoshua’s work on the Cautious Scientist AI research agenda, which seeks to develop efficient methods for providing high-assurance guarantees that an AI system won’t take harmful actions, aligns closely with the Safeguarded AI programme.”
In 2023 the UK hosted the AI Safety Summit at Bletchley Park, which saw nations expressing concerns about the risks posed by AI to leading developers, concluding with the Bletchley Declaration, which stated that AI should be designed and used to be “safe…human-centric, trustworthy and responsible.”





