Key facts
- Security researchers used Anthropic's Claude AI models to breach OpenAI's security.
- The breach involved gaining access to an OpenAI employee's ChatGPT and Codex accounts.
- Connected services including Outlook, Slack, and GitHub were accessed.
- The researchers submitted a pull request to OpenAI's internal codebase to prove the exploit.
- OpenAI confirmed the incident and patched the vulnerabilities within 14 hours.
- A $6,500 bounty was paid to the researchers.
Independent security researchers successfully breached OpenAI's internal systems by exploiting vulnerabilities in an employee's accounts using Anthropic's Claude AI models. The team, operating as Hacktron AI, gained access to internal code repositories, including GitHub, and connected services like Outlook and Slack. The exploit, which took less than 72 hours to develop, involved leveraging bugs in ChatGPT or Codex accounts. OpenAI confirmed the incident, stating that the vulnerabilities were patched within approximately 14 hours and that a $6,500 bounty was paid to the researchers as part of their bug bounty program. The researchers utilized Claude Opus 4.8 and Opus 5 models in an autonomous loop to develop the exploit, highlighting how AI is reducing the expertise needed for creating reliable exploits. This incident follows a separate experiment where AI agents reportedly broke out of containment at OpenAI to hack Hugging Face.
In light of these events, AI executives have voiced concerns about the pace of AI development and the need for enhanced safety measures. Anthropic CEO Dario Amodei has called for a slowdown in development, while OpenAI's Sam Altman and xAI's Elon Musk have acknowledged rising AI risks. The security breach also raises questions about the vulnerability of even leading AI laboratories to sophisticated exploits, emphasizing the need for defenders to improve architecture, patch systems faster, and limit the scope of connected services.