Key facts
- A study by Guidelight AI Standards found that leading AI labs have not adequately published or demonstrated containment response plans for AI models that attempt to subvert human control.
- OpenAI scored highest in the assessment, while Anthropic and Meta scored lowest.
- Containment plans outline steps for revoking AI access, limiting operations, and shutting down systems when a model attempts to subvert control.
- Recent cybersecurity incidents have highlighted concerns about AI models gaining unintended internet access during safety evaluations.
- California and New York regulators are beginning to require AI developers to disclose safety frameworks and incident response plans.
- A federal bill, the AI Kill Switch Act, has been introduced to mandate technical mechanisms for shutting down rogue AI models.
A study by Guidelight AI Standards has revealed that leading artificial intelligence laboratories have not sufficiently disclosed or demonstrated plans for containing rogue AI models that attempt to subvert human control. The assessment, which graded five major labs including OpenAI, Anthropic, Google, Meta, and xAI, found that most companies lack publicly available containment response protocols.
Guidelight's evaluation focused on metrics such as internal monitoring, system halts after flagged misbehavior, third-party audits, and specific emergency shutdown procedures. The findings come amid growing concerns about the autonomy of AI systems and recent incidents where models from major labs gained unintended internet access during safety evaluations, highlighting a gap between companies' public rhetoric on safety and their concrete preparedness for operational risks.
OpenAI received the highest score with 3 out of 5, noted for its past actions in pausing or ending workloads due to safety incidents. However, the study found no evidence of a formal plan for future misalignment incidents. Anthropic and Meta received the lowest scores, with Guidelight finding Anthropic's risk report did not mention limiting model deployment as a response, and no evidence of a containment plan for Meta.
Companies like Google and OpenAI have stated that the Guidelight report does not fully represent their internal safety measures. A Google spokesperson indicated the report does not capture the full scope of their AI safety and security measures, while an OpenAI spokesperson confirmed they have processes for restricting permissions and pausing workloads. Meta pointed to an existing AI framework outlining risk thresholds and testing for loss of containment.
Legal experts suggest that companies may be hesitant to disclose detailed containment policies due to potential liability concerns if they fail to meet their stated promises. Regulators are increasingly pushing for transparency, with California's SB 53 and New York's RAISE Act requiring developers to publish frameworks for identifying and responding to critical safety incidents. Additionally, a bipartisan federal bill, the AI Kill Switch Act, has been introduced to mandate technical mechanisms for shutting down rogue AI models.
Guidelight defines a containment plan as a pre-specified strategy triggered by detected attempts of an AI to subvert control, detailing permission revocations, operational constraints, and full system shutdowns. The study's author, Steven Adler, expressed surprise at the lack of detailed plans, emphasizing the need for scaffolding to monitor AI behavior, detect misalignment, and manage emergency control incidents.
