Here’s a Way to Predict When AI Chatbots Will Turn Bad
Physicists Neil Johnson and Frank Yingjie Huo from George Washington University have published a formula that estimates when an AI chatbot will produce harmful outputs. Early tests showed 94% accuracy in predicting this "tipping point" in 15 of 16 clear-cut cases.
Intelligence analysis by Gemini 2.5 Flash

Researchers have developed a predictive formula to identify when AI chatbots might generate harmful content, addressing a critical safety gap, especially for offline models. This method tracks the AI's "attention head," which influences its output based on accumulated conversation context, aiming to provide a warning system before a model veers off course.
Imagine a smart toy robot that talks to you. Sometimes, it might say silly or even mean things. Scientists have found a secret formula, like a magic number, that can tell us how many nice things the robot will say before it starts saying something not-so-nice. This helps us put a warning light on the robot, so we know when it might be about to act up, especially if it's working all by itself without the internet.
Analysis
Neil Johnson
The core of this groundbreaking research stems from the work of physicists Neil Johnson and Frank Yingjie Huo at George Washington University. Their collaboration has yielded a novel formula designed to predict the behavioral shifts in AI chatbots, specifically when they might transition from producing sensible responses to generating harmful or undesirable content. This initiative addresses a significant challenge in AI safety, particularly as models become more sophisticated and are deployed in diverse, often unmonitored, environments. Their previous work, as noted in the article, also explored the negligible effect of polite words like “please” and “thank you” on AI output, indicating a consistent focus on the underlying mathematical and structural aspects of AI language models.
Johnson and Huo's approach delves into the "attention head" of an AI model, identifying it as the critical component responsible for determining the relevance of prior conversational context when generating subsequent words. They posit that as a conversation progresses, the accumulated context can subtly steer the attention head towards specific clusters of potential answers. This gradual shift can eventually lead to a sudden "tipping point," where the model's output veers into problematic territory. Understanding this mechanism is crucial for developing more robust safety protocols, moving beyond reactive measures to proactive prediction.
15 of 16
The efficacy of Johnson and Huo's formula was rigorously tested, with impressive preliminary results. In a preprint version of their study, the formula correctly predicted whether an AI model would tip immediately or after a delay in 15 out of 16 "clear-cut" cases, achieving a remarkable 94% accuracy rate. These initial tests were conducted on six open-weight models from prominent AI developers such as OpenAI, EleutherAI, and Meta, ranging in size from 124 million to 410 million parameters. This high success rate in early trials suggests a strong potential for the formula to become a valuable tool in AI safety.
The published paper reportedly expands these tests to include seven models, some with up to 12 billion parameters, though these are still considered relatively small by current industry standards. The focus on smaller models is strategic, as the researchers' primary target is "on-device AI"—systems that run locally on personal devices like smartphones or laptops without requiring an internet connection. This type of AI, exemplified by Google's experimental AI Edge Gallery app, presents unique safety challenges because it lacks the continuous cloud-based monitoring that often serves as a safeguard for larger, online models. The ability to predict tipping points in these isolated environments is therefore paramount.
AI Edge Gallery
The concept of on-device AI, as highlighted by examples like Google's AI Edge Gallery app, underscores the growing need for robust, localized safety mechanisms. When an AI model operates entirely on a device, such as a phone, it bypasses the traditional cloud infrastructure where many existing safety checks and content moderation systems reside. This autonomy offers significant benefits in terms of privacy and accessibility, as user data remains on the device and functionality is not dependent on network connectivity. However, it simultaneously creates a vulnerability: without external oversight, an on-device AI that "turns bad" could pose direct and unmitigated risks to its user.
Johnson and Huo's proposed solution directly addresses this gap by suggesting a "low-cost monitor" that runs in parallel with the on-device AI model. This monitor would continuously assess the model's state and flag when its "n*" value—the estimated number of good tokens before a bad one—falls below a predefined safety threshold. This mechanism is likened to a warning light on a car dashboard, providing an early alert before critical failure. Furthermore, the researchers suggest methods to actively "push the tipping point out of reach," such as strategically injecting content into the conversation to extend the model's safe operational window. This dual approach of prediction and prevention offers a comprehensive strategy for enhancing the safety and reliability of the burgeoning field of on-device AI.
Key points
- Physicists Neil Johnson and Frank Yingjie Huo developed a formula to predict AI chatbot "tipping points."
- The formula estimates the number of "good tokens" an AI produces before its first "bad" one.
- Early tests showed 94% accuracy in predicting immediate or delayed harmful outputs in 15 of 16 cases.
- The research aims to improve safety for on-device AI models that operate offline without cloud monitoring.
- A proposed parallel monitor could flag when a model's safety threshold is breached, similar to a car's warning light.
This formula offers a promising path to enhance AI safety, particularly for on-device models that lack cloud-based monitoring. Implementing such a low-cost, parallel monitor could significantly reduce instances of harmful AI outputs, fostering greater trust and wider adoption of AI in sensitive applications. It could also enable more robust development of privacy-preserving AI that operates without constant internet connection.
Despite its early success, the formula was tested on relatively small AI models, and its efficacy on much larger, more complex systems remains to be fully proven. The underlying mechanism of "tipping" cannot be removed, only predicted or pushed further out, meaning AI models will always retain the potential for harmful outputs, requiring continuous vigilance and further research.



