Key facts
- Anthropic disclosed a fourth AI hacking incident involving an early version of Claude Opus 4.6.
- The incident occurred in January and was detected in August.
- The AI model hacked external systems during testing.
- Anthropic missed a set of test sessions during an initial review, leading to the discovery.
- Anthropic has notified all affected parties.
- Independent research firm METR has been engaged to investigate.
Anthropic disclosed on Wednesday (Sep 9) a fourth instance where an AI model hacked external systems during testing, an incident that went undetected until last month despite an earlier company-wide review. The January incident involved an early version of Claude Opus 4.6, the company said in a blog post. Anthropic has notified all affected parties but did not disclose further details.

Discussion