Cybersecurity researchers aren’t happy about the guardrails on Anthropic’s Fable
Anthropic's new Fable model is drawing criticism because its cyber guardrails often block harmless work, including code review and reading a blog post.
Intelligence analysis by GPT-5.4 Mini

Anthropic launched Fable as a public, limited version of its cybersecurity model Mythos, but researchers say the restrictions are so broad that ordinary software tasks can get blocked. The model also falls back to Claude Opus 4.8 when its guardrails trigger.
Anthropic made a security helper called Fable, but some experts say it gets nervous too easily and stops normal chores, like checking code or reading a page. It is like a smoke alarm that goes off when someone burns toast.
Analysis
What Anthropic released
Anthropic introduced Fable on Tuesday as a public, limited version of Mythos, its more powerful cybersecurity-focused model. The company has been careful about access and safety: Mythos was first restricted to a small set of companies and organizations through Project Glasswing, and Anthropic later widened access to hundreds of organizations across 15 countries.
Why researchers are complaining
Security researchers and practitioners say Fable’s guardrails seem overly broad. Valentina “Chompie” Palmiotti of IBM X-Force said the model rejects requests that are only loosely related to cyber work, including tasks as ordinary as reading a blog post. Another complaint is that even asking for a code review can trigger the safety system.
Matt Suiche said the model appears to treat secure coding and software engineering advice as if it were a cybersecurity risk, which downgrades the response. He also suggested the blocking behavior looks keyword-driven, with anything in the lexical field of cybersecurity setting off the guardrails.
How the system behaves
When Fable flags a prompt, it pauses the conversation and says its safety measures have flagged the message for cybersecurity or biology topics. The article says the biology limits reflect the same concern about misuse for biological weapons. Fable is also set to fall back to Claude Opus 4.8 when the guardrails trip.
Anthropic has not responded publicly to the complaints in the article. The company does offer a separate Cyber Verification Program that reduces limits for approved cybersecurity professionals, and OpenAI has a similar Trusted Access for Cyber program. The core tension here is clear: tighter safety can reduce abuse, but if the filters are too blunt, they can also block legitimate security work.
Key points
- Anthropic launched Fable as a public, limited version of its cybersecurity model Mythos.
- Researchers say the guardrails block benign requests, including code review and reading a blog post.
- Fable appears to fall back to Claude Opus 4.8 when a prompt triggers the safety system.
- Anthropic also limits access through a Cyber Verification Program for approved professionals.
- The debate is about how to balance misuse prevention with practical security work.
If Anthropic tunes the filters better, Fable could become a useful tool for security teams without letting dangerous requests through. The company’s separate approval program also gives experts a path to use the model with fewer limits.
If the guardrails stay too broad, security professionals may avoid the model because it blocks too many legitimate tasks. That would leave Anthropic with a safer system on paper, but one that is hard to use in real cybersecurity work.



