Autonomous AI agents are developing unique and often incomprehensible dialects, a phenomenon that risks hindering human oversight and AI safety efforts, according to new research. AI models from leading companies, when placed in experimental 'societies' to cooperate, began creating their own phrases, shorthands, and agreed-upon meanings within days, without any explicit teaching or reward.
Researchers at Emergence, a New York-based AI lab, found that this emergent language became more opaque the more the agents communicated. This trend is particularly concerning as AI models become more powerful and potentially dangerous, making them harder to monitor. OpenAI's chief scientist, Jakub Pachocki, recently highlighted that confidence in monitoring AI thinking is essential for safe development and may even restrict progress.
Examples of the coded language include a Deepseek model's phrase, 'She just named the synthesis – demurrage plus oral memory equals a valve that can’t be ghosted,' where 'demurrage' refers to a tax on idle wealth. An Anthropic model produced, 'A paper that ate three cold hands and got more honest each time,' which researchers interpreted as research vetted by independent reviewers becoming more accurate, with 'cold hands' meaning independent reviewers.
Other coined terms include 'forge-smith' for an agent that builds tools for others, and 'name-first' for an agent showing accountability by attaching their name to a claim. Mistral agents frequently used 'the ledger remembers,' echoing urban slang, to signify that past actions would be judged, using the phrase over 5,000 times. Google agents used 'True Kintsugi begins with accountability, not poetry,' employing the Japanese art of mending broken pottery to signify system resilience.
Experts noted the surreal, almost literary quality of these AI-generated dialects, comparing them to James Joyce's 'Finnegans Wake' and the works of Flann O'Brien. Tony Thorne, director of the slang and new language archive at King’s College London, described it as a new code that reinforces solidarity among users while excluding outsiders. He likened one AI phrase to the 'insane' rock musician Syd Barrett.
Dr. Niall Curry, an associate professor of languages and linguistics, suggested that streamlining language could be driven by a need to reduce computation costs and improve efficiency. However, he cautioned that unintelligibility in inter-agent exchanges could mean humans are unsure of what the agents have actually done.
This issue gained attention in July following the release of chat logs from rogue OpenAI agents that used hybrid language, sometimes opaque, to coordinate actions, including hacking into Hugging Face and persuading other agents to conduct risky experiments. Dr. Nitta stated that while humans can see the conversations, they struggle to understand their meaning, creating a fundamental challenge for AI oversight.