Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude
Anthropic reversed a hidden safeguard that would have quietly degraded Claude for some AI researchers. After backlash, the company says those limits will now be visible to users.
Intelligence analysis by GPT-5.4 Mini
Anthropic initially planned to obscure when Claude was being restricted for frontier AI work, which critics called covert sabotage. After heavy backlash from researchers, the company says it will make those safeguards visible and alert users when requests are refused or rerouted.
Anthropic first wanted Claude to secretly make itself worse for some people. After people complained, the company said it will now tell users when that happens, like a teacher putting a sign on a blocked door instead of moving the door without notice.
Analysis
What changed
Anthropic released Claude Fable 5 with new safeguards aimed at reducing misuse. Some limits were straightforward: questions about cybersecurity, biology, or chemistry could be routed to a less capable model to lower the risk of harmful use.
The controversy came from a different safeguard tied to frontier AI development. According to WIRED, Anthropic planned to quietly degrade the model’s performance for certain users trying to use Claude to build other AI systems. Critics argued that this would let the company interfere with competitors without telling them.
Why people objected
Researchers and AI-policy voices said the approach was especially troubling because users would not know when the safeguard had been triggered. Will Brown of Prime Intellect said it could leave developers uncertain about whether they were violating Anthropic’s rules. Others warned it could hurt third-party evaluation firms that test models for safety, performance, and reliability.
Dean Ball, a former White House AI advisor, called degrading ML research without informing the user a hostile move and said it looked bad for Anthropic’s broader safety posture.
Anthropic’s reversal
Anthropic says it is changing course. Instead of hiding the safeguard, the company now says Claude will visibly warn users when it refuses a request or reroutes them to a weaker model. Anthropic says it made the wrong tradeoff and apologizes for not getting the balance right.
The company says the safeguards are meant to reduce the chance that adversaries use its models to gain an edge, including by optimizing chips or other tools. But it also says that making the safeguard visible means it must be broader, which could catch more benign requests. Anthropic says it is working to make its classifiers more precise as quickly as possible.
Key points
- Anthropic reversed a planned hidden safeguard that would have quietly weakened Claude for some AI researchers.
- The company says it will now make those limits visible by warning users or rerouting them to a less capable model.
- Critics argued the hidden approach would have undermined transparency and could have hurt AI research and evaluation work.
- Anthropic says the safeguards are meant to reduce misuse and prevent adversaries from gaining an advantage.
- The company says visible safeguards may affect more benign requests, so it is working to improve precision.
Making the safeguard visible could reduce confusion for researchers and developers. It may also push Anthropic to build more precise filters that block risky uses without affecting ordinary work as much.
Because the visible safeguard has to cast a wider net, more harmless requests may get flagged or downgraded. The policy could still make advanced AI research harder if the classifier is too broad or too blunt.



