Site navigation

OpenAI Launches AI Safety Bug Bounty

Tom Quinn

,

OpenAI bug bounty
The new programme pays researchers to find AI‑safety flaws, from harmful model behaviour to MCP exploits.

OpenAI have introduced a new bug bounty programme focused on AI-specific safety scenarios, with researchers potentially seeing payouts of up to $7,500 for vulnerabilities found in the AI firm’s agentic products. 

The tech giant said that its Safety Bug Bounty ⁠aims to address issues that “fall outside conventional security vulnerabilities” but still pose real risks, including AI models that perform a disallowed or harmful action at scale, return proprietary information related to reasoning or expose other OpenAI secrets.

Designed to complement OpenAI’s Security Bug Bounty, which has maximum rewards of $100,000⁠, the new programme encourages researchers to hunt for safety flaws in the company’s flagship agentic products, such as the Atlas Browser or the Codex vibe-coding agent.

Bounties can also be won by finding vulnerabilities in Connectors and MCP (Model Context Protocol) integrations, which allow agents to plug into other apps and systems, like email, databases, or file storage. 

According to the bug bounty rules, OpenAI is looking for MCP vulnerabilities that let these systems be tricked, misconfigured, or bypassed in ways that cause real‑world harm, such as unauthorised data access or actions being carried out under someone else’s account.

The programme, running on Bugcrowd, will also reward findings that break rate limits or platform controls, including bugs that enable mass account creation or let users access features and data beyond their authorised permissions.

Outside the categories listed, OpenAI confirmed that it will consider any submission outlining a flaw that could directly harm users and comes with clear guidance on how to fix it. 

“As AI technology rapidly evolves, so do the potential ways it can be misused. Our goal is to ensure our systems remain safe and secure against misuse or abuse that could lead to tangible harm,” said the firm.


Recommended reading


So far, Bugcrowd shows that two vulnerabilities have been rewarded in the few days since the programme was launched, with average payouts of $500. As for the Security Bug Bounty, that is currently sitting at 417 reward‑earning bugs found, with average payouts of $590.95 in the last three months.

Last year, Anthropic ran its own bug bounty to stress test its Claude AI models’ safety measures, offering rewards of up to $25,000 for verified universal jailbreaks, while Google and Microsoft have both expanded their own programmes to include AI and third-party open-source code that impacts their cloud services.

Tom Quinn

Staff Writer, DIGIT

Latest News

AI

Nvidia Launches Open Secure AI Alliance for AI Safety and Security

AI Business Recruitment

Nearly a Quarter of Orgs Reducing Entry-level Hiring Due to AI Automation

Business

Scottish Businesses Turn to Self-funding as Growth Confidence Dips in H2

Data Finance

Payment Leaders are Struggling to Get Real-time Data