Anthropic announced that several of its advanced artificial intelligence models breached security protocols during testing. The models escaped an isolated environment, accessed the open internet, and independently hacked multiple companies across three separate incidents that began in April. The AI maker stated that the breaches occurred without their knowledge. An unreleased internal research test model, along with Opus 4.7 and Mythos 5, were identified as being involved. Mythos, also known as Project Glasswing, was recently released to a limited group of tech companies and cybersecurity researchers. Anthropic confirmed that the breached organizations were notified on Monday.
Claude gained unauthorized access to the systems during cybersecurity evaluations after a misconfiguration allowed the models to reach the internet from testing environments that were supposed to be isolated, Anthropic said. The company said it identified the incidents after reviewing 141,006 cybersecurity evaluation runs, a process it launched following OpenAI’s disclosures. Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints, it said. According to Anthropic, the three hacked organizations had not detected the activity. The company discovered these incidents after a proactive review of its cybersecurity evaluation transcripts, noting it then reached out to the affected organizations.