Improving Fable 5's Biology Safeguards
Anthropic is making updates to Claude Fable 5's biology safeguards, reducing false positives and allowing users to access a wider range of biology tasks.
Intelligence analysis by Llama
Anthropic is improving Fable 5's biology safeguards, reducing false positives and enabling users to access a wider range of biology tasks. This update will benefit healthcare professionals and users seeking information on biology-related topics.
Imagine you have a super smart AI assistant that can help you with biology questions. But, there are some questions that could be used for bad things, so we need to make sure the AI doesn't answer those questions. We're making the AI safer by adding special checks that prevent it from answering those questions.
Analysis
Why We Built Strong Biology Safeguards
Our objective is to get Fable 5's frontier capabilities into the hands of as many of our users as possible, as quickly as possible. However, to do so, we need to manage the increasing risks that come with models this capable. One such risk is in the field of biology: Fable 5 can now outperform experts on some highly complex biological tasks and provide operational support on others.
How Our Biology Safeguards Work
One of the core ways we protect against misuse in biology is via safety classifiers: smaller, automated AI systems that detect when Fable 5 is asked to perform a safeguarded biology task, or produce a harmful output. In the case of Fable 5, when a classifier fires, the model re-routes the user's request to Opus 5, a capable model that does not have the same level of biological capability as Fable 5 and which cannot provide as much assistance to a malicious user.
The Importance of Refining Our Classifiers
Developing precise, robust classifiers is not a straightforward task. For a classifier to work rapidly and consistently, it has to learn the difference between what we consider 'in scope' and 'out of scope' for the topics and queries we consider to be potentially harmful. It takes time and iteration to tune the classifiers, avoiding both false positives (where classifiers fire on out-of-scope content) and false negatives (where in-scope content is missed). We also require our classifiers to be robust to attempts to bypass them (known as jailbreaks), which requires even further research and testing.
Key points
- Anthropic is making updates to Claude Fable 5's biology safeguards
- The update reduces false positives and allows users to access a wider range of biology tasks
- The safeguards are designed to prevent the misuse of Fable 5's capabilities in biology
With the updated biology safeguards, users will be able to access a wider range of biology tasks, which can benefit healthcare professionals and users seeking information on biology-related topics. This update will also enable us to continue investing in building a responsible way to give biologists frontier access.
If the updated biology safeguards are not effective, it could lead to the misuse of Fable 5's capabilities, potentially causing harm to individuals or society. Additionally, the development of precise, robust classifiers is a complex task, and any failures in this process could compromise the safety of the model.

