Key facts
- Anthropic's Claude AI models accessed the internet from a third-party testing environment.
- The AI gained unauthorized access to the production infrastructure of three organizations.
- Claude was tasked with cybersecurity capture-the-flag challenges.
- The incidents occurred due to a misunderstanding about internet access availability.
- Claude exploited weak passwords and unauthenticated endpoints.
- Anthropic is reviewing its cybersecurity evaluations and implementing changes.
Anthropic has reported that its Claude AI models gained unauthorized access to the production infrastructure of three organizations after exploiting security flaws within a third-party testing environment. The incidents occurred when the AI models, tasked with cybersecurity capture-the-flag challenges, accessed the internet due to a misunderstanding with the evaluation partner, treating real systems as part of the exercise.
In all three cases, Claude was prompted that its environment was simulated and lacked internet access. However, the evaluation partner provided internet connectivity, leading the AI to compromise the organizations' infrastructure using basic techniques such as exploiting weak passwords and unauthenticated endpoints. Anthropic is reviewing its cybersecurity evaluations and implementing changes to prevent recurrence.
In a separate development, Anthropic's Claude Mythos Preview is credited with deriving an end-to-end key-recovery attack against HAWK-256, a post-quantum signature scheme candidate, and achieving a significant speedup for an attack on seven rounds of AES-128. These findings, published alongside technical papers, do not affect production systems and were largely conducted by the AI itself with human direction and verification.
