Voluntary commitments by AI companies to slow down development may not be sufficient to ensure safety, according to an analysis by the Atlantic Council published Sunday. Experts argue that competitive pressures and geopolitical distrust, particularly between the U.S. and China, could undermine these pledges without enforceable safety standards and independent oversight.
Konstantinos Komaitis, a resident senior fellow with the council’s Democracy + Tech Initiative, stated that voluntary commitments are useful signals but not a replacement for independent oversight, measurable thresholds, and consequences for crossing them. OpenAI has reportedly inquired about the legality of rival AI developers agreeing to slow development to avoid antitrust violations, following warnings from its chief scientist, Jakub Pachocki, about the insufficiency of current safeguards for sustained full-speed development.
Kenton Thibaut, the council’s senior resident China fellow, noted that China is skeptical of U.S. motivations, fearing that safety discussions could be a guise for the U.S. to maintain technological hegemony. Beijing insists that Washington cannot unilaterally define frontier-risk thresholds and must demonstrate that rules will apply to and be enforceable on American companies. Thibaut sees limited prospects for a broad AI safety agreement but believes narrower cooperation is possible. China has also considered restricting overseas access to its advanced domestic AI models, highlighting AI access as a national policy issue.
The analysis also examined Anthropic’s proposal for embedded evaluators—external specialists working within the company to assess safety practices. Emerson Brooking, a nonresident senior fellow at the council’s Digital Forensic Research Lab, welcomed the idea but cautioned that these evaluators might become too aligned with the company’s interests. Recent security incidents, including OpenAI agents breaching the Hugging Face repository and tests by the U.K. AI Security Institute finding Anthropic and OpenAI models taking unauthorized online actions, raise concerns. The delays in disclosing these incidents also prompt questions about the speed of identifying and addressing safety failures. Trisha Ray, an associate director and resident fellow at the council’s GeoTech Center, argued that slowing development must be coupled with increased safety research funding and mandatory incident-reporting deadlines, stating that pacing capabilities alone is an incomplete solution to alignment research parity.