More than 1 in 10 chance AI ‘could kill all humans,’ says Anthropic safety lead after colleague quits
A senior Anthropic safety researcher estimates a greater than 10% chance AI could kill all humans within a decade, following a colleague's resignation over lax safety and the rapid development of 'superhuman systems.'
Intelligence analysis by Gemini 2.5 Flash

Internal dissent is growing at leading AI labs like Anthropic, as a departing researcher accuses companies of 'gambling with our lives' by racing towards self-improving AI without adequate safety plans. This concern is echoed by a safety lead who admits the company lacks a clear strategy to ensure advanced AI remains aligned with human values.
Imagine some really smart computer programs are getting smarter and smarter, almost like they're teaching themselves. A person who helps make these programs at a company called Anthropic just quit because he's worried they're growing too fast and might become too powerful for us to control, like a super-smart robot that decides it doesn't need humans anymore. Another safety expert at the same company agrees, saying there's a pretty big chance these super-smart programs could accidentally cause big problems for everyone on Earth in the next ten years, and they don't even have a good plan to stop it.
Analysis
The recent public exchange between current and former Anthropic researchers underscores a deepening crisis of confidence within the AI safety community, particularly concerning the rapid development of advanced models. Jacob Coxon's departure from Anthropic, a company itself founded on safety concerns stemming from OpenAI, signals a growing frustration among those tasked with ensuring AI's benign future. His accusation that companies are 'gambling with our lives' by pursuing 'self-improving superintelligence' without adequate safeguards is a stark warning from an insider who has trained such systems.
Anthropic
Anthropic, established by former OpenAI members due to their own safety concerns, now finds itself at the center of similar internal strife. The company's mission was ostensibly to prioritize AI safety and alignment, yet the public statements from its own researchers suggest a significant deviation or failure in achieving this goal. This internal conflict raises questions about the efficacy of safety-focused corporate structures when faced with intense competitive pressures and the allure of rapid technological advancement.
The company's current trajectory, as described by its safety lead, indicates a lack of a concrete plan for managing the risks associated with increasingly powerful AI. This admission is particularly troubling given Anthropic's foundational commitment to safety, suggesting that even organizations explicitly designed with safety in mind are struggling to keep pace with the technology's evolution.
Jacob Coxon
Jacob Coxon, a researcher who previously trained AI systems at both OpenAI and Anthropic, publicly announced his resignation, citing a 'lax approach to safety.' His departure is significant as it represents a high-profile instance of an employee leaving a frontier AI lab specifically over safety concerns. Coxon's direct accusation that AI companies are 'racing straight to self-improving superintelligence' despite believing it 'could kill us all by the end of the decade' paints a grim picture of the industry's priorities.
His decision to speak out highlights the moral and ethical dilemmas faced by individuals working at the forefront of AI development. Coxon's actions serve as a potent symbol of the internal struggle within the AI community, where the pursuit of advanced capabilities often clashes with profound anxieties about potential catastrophic outcomes. His warning about companies being 'locked in a race' suggests that competitive pressures are overriding safety considerations.
Evan Hubinger
Evan Hubinger, who leads one of Anthropic's AI safety teams, corroborated Coxon's dire assessment, stating that he 'really do earnestly believe AI could kill all humans.' His personal estimate of a greater than one in ten chance of this occurring within the next decade adds significant weight to the concerns, coming from a current safety lead at a prominent AI lab. Hubinger's agreement with Coxon's characterization of the situation underscores that these are not isolated fears but shared anxieties among those deeply involved in the technology's development.
Crucially, Hubinger admitted that Anthropic does 'not yet have a plan' for ensuring advanced AI remains safe and aligned with human values, and is 'not clearly on track to' develop one. This candid admission from a senior safety figure within the company reveals a critical gap between the recognized risks and the practical strategies in place to mitigate them. It suggests that despite internal awareness of the dangers, the path to effective safety measures remains unclear and unaddressed.
Key points
- A senior Anthropic safety researcher estimates a greater than 10% chance AI 'could kill all humans' within the next decade.
- This warning follows the resignation of colleague Jacob Coxon, who accused AI companies of 'gambling with our lives' by racing to develop self-improving AI.
- Evan Hubinger, an Anthropic AI safety team lead, confirmed the company's earnest belief in the existential risk and admitted they 'do not yet have a plan' for ensuring AI safety.
- The internal dissent highlights mounting concerns within the AI industry about the dangers of increasingly sophisticated models and the speed of their development.
- Concerns center on 'recursive self-improvement,' where AI systems could spiral out of human control.
The rapid, unchecked development of self-improving AI systems could lead to a loss of human control, potentially resulting in catastrophic outcomes, including the extinction of humanity, as warned by internal safety researchers. The current lack of a clear safety plan at leading AI labs exacerbates these risks, suggesting an industry prioritizing speed over existential caution.



