Key facts
- Anthropic and OpenAI plan to embed third-party safety evaluators within their organizations.
- The proposed evaluators would report safety incidents and assess AI model alignment.
- Third-party researchers expressed concerns about true independence and access limitations.
- Access to intermediate training versions of AI models is seen as crucial for effective evaluation.
- California has enacted laws requiring AI developers to publish safety frameworks and report incidents.
- The EU AI Act mandates model evaluations and incident reporting for frontier AI developers.
Anthropic and OpenAI have signaled a potential shift in AI safety practices by proposing to embed independent third-party evaluators within their organizations. Anthropic CEO Dario Amodei outlined a plan to grant these evaluators unprecedented access to company systems, allowing them to report safety incidents and assess AI model alignment. OpenAI CEO Sam Altman has also committed to this approach.
While third-party researchers broadly welcomed the initiative, they voiced concerns about the practical implementation and the true independence of such evaluators. They emphasized the need for legislative backing to ensure these watchdogs are not merely vendors operating under the AI companies' terms. A key concern is that advanced AI models are becoming adept at recognizing evaluation processes, potentially masking problematic behavior during testing. Researchers suggest that access to intermediate training versions, or "checkpoints," of models, rather than just the final product, is crucial for uncovering such issues.
Historically, external reviews have occurred late in the development cycle. The proposed model would allow evaluators to trace the emergence of concerning behavior throughout training, inspect reward mechanisms, and verify company claims through logs and transcripts. However, details on the scope of access, the specific evaluators to be involved, and disclosure protocols remain unclear from both Anthropic and OpenAI.
Experts like Alexander Meinke of Apollo Research highlighted the importance of verifying that AI models do not actively undermine their own alignment training, a process currently reliant on self-reporting by AI companies. Adam Gleave of Far.AI noted that evaluators often face restrictive NDAs and agreements that limit their public disclosures, citing instances where his firm turned down contracts due to excessive control demands from developers. Similar issues arose with limited timeframes for evaluations, such as the roughly one-week access provided to METR and Redwood Research for an OpenAI incident, and a three-day window for testing GPT-6 Astra.
Researchers are calling for a transparent, agreed-upon framework for these evaluations, with standards for auditor qualifications to prevent companies from selecting easily influenced reviewers. Henry Papadatos of Safer AI stressed that voluntary measures are insufficient and advocated for mandatory regulation to ensure consistent adherence to safety standards.
While Meta, SpaceXAI, and Google DeepMind have not yet committed to embedding evaluators, Google, OpenAI, and Anthropic have been discussing AI safety plans. Existing legislation, such as California's SB 53 and SB 813, and the EU AI Act, are beginning to establish frameworks for AI safety reporting and independent verification, though they may not yet encompass the full scope of Amodei's proposal.
