Anthropic Releases Claude Fable 5, Its Most Powerful AI Yet, With Cyber Safeguards
Anthropic launched Claude Fable 5 for the public and kept a more powerful twin, Claude Mythos 5, restricted to vetted cyber defenders.
Intelligence analysis by GPT-5.4 Mini

Anthropic split its newest model into two public-facing products: one with cyber safety filters for general users, and one with those safeguards relaxed for approved defenders and critical infrastructure teams. The release highlights how capable frontier models have become at both finding bugs and preventing abuse.
Anthropic made a very smart robot brain and then gave the public a safer version while keeping the sharper tools locked away for trusted security experts. It is like handing out kitchen scissors to everyone but storing the sharp chef’s knife behind a locked door.
Analysis
What Anthropic shipped
Anthropic says Claude Fable 5 is its most capable model to date and is now generally available. The company also introduced Claude Mythos 5, which uses the same underlying model but keeps cyber capabilities available only to a vetted group of cyber defenders and critical infrastructure operators.
How the split works
The public Fable 5 does not simply refuse risky requests. Instead, when a request is flagged for cyber, biology, chemistry, or model-distillation concerns, it is routed to Claude Opus 4.8, a weaker fallback model. Anthropic says the cyber classifier is intended to block offensive tasks such as reconnaissance, vulnerability discovery, lateral movement, and other steps that could support a real attack. Users are told when that handoff happens.
The company says the design is intentionally conservative. That means it can catch harmless requests too, and Anthropic says fallback happens in under 5% of sessions. It also says it plans to reduce false positives after launch.
What the safety testing showed
Anthropic says external testing did not find a universal jailbreak, even after more than 1,000 hours of bug bounty work. It also says outside red teams did not break the safeguards during long agentic tasks, though the UK AI Security Institute reportedly made progress toward a universal jailbreak in an early test window. Anthropic says the goal is not perfect prevention, but making abuse slow and expensive enough to detect before it scales.
Why the release matters
The article frames the product split as a response to a real capability problem. Anthropic says the underlying model family can already do serious offensive work if unrestricted, so the company is trying to give the public a safer version while preserving stronger capabilities for vetted defenders. That tradeoff will likely become a template for how frontier AI systems are governed in security-sensitive domains.
Key points
- Anthropic released Claude Fable 5 to the public and kept Claude Mythos 5 restricted to vetted cyber defenders.
- The public model routes flagged cyber and related requests to the weaker Claude Opus 4.8.
- Anthropic says the safeguards blocked harmful cyber requests in testing, but can produce false positives.
- External red teaming and bug bounty work did not uncover a universal jailbreak, though one government lab made some progress.
- The company argues the model’s offensive capability is real enough that access controls are necessary.
If the safeguards work as intended, more people get access to a powerful model without turning it into an easy cyberattack tool. Security teams still keep access to stronger capabilities, which could help them find bugs and defend important systems faster.
The article makes clear that the safeguards are not perfect and can still block harmless requests, which could frustrate legitimate users. Anthropic also acknowledges that universal jailbreaks may be impossible to stop completely, so determined attackers may still find ways around the controls.



