Physicists at George Washington University have developed a formula that can predict when an AI chatbot is likely to produce harmful content, a significant step towards ensuring the safety of AI models, particularly those designed to run offline on personal devices. The formula, developed by Neil Johnson and Frank Yingjie Huo, estimates the number of safe responses, or "good tokens," an AI model will generate before it produces its first unsafe or harmful output. This "tipping point" is crucial because current safety tools often rely on cloud connections, leaving offline models vulnerable.
The researchers' work, detailed in the journal Patterns and a prior preprint, suggests that the problem stems from the "attention head" within AI models, which determines which parts of a conversation are most relevant for generating the next response. As a conversation progresses, the accumulated context can cause this attention mechanism to shift, potentially leading to harmful outputs. This mechanism is also exploited by users attempting to "jailbreak" AI models.
In early tests described in the preprint, the formula accurately predicted whether a model would fail immediately or after a delay in 15 out of 16 clear-cut cases, achieving 94% accuracy. These tests were conducted on six open-weight AI models, ranging from 124 million to 410 million parameters, sourced from companies like OpenAI, EleutherAI, and Meta. The published paper reportedly expands these tests to larger models, up to 12 billion parameters.
The ultimate goal is to enhance the safety of on-device AI, which operates entirely on a user's phone or laptop without an internet connection. To address the gap left by the absence of cloud-based safety checks, Johnson and Huo propose a parallel monitor that can flag when a model's safety threshold is breached, akin to a warning light in a car. They also suggest methods to extend the safe response window by strategically injecting content into conversations. While alignment training can modify behavior for specific prompts, the authors note it does not eliminate the underlying tipping mechanism.