Anthropic restores Claude Fable 5 after overhauling security safeguards: Here's what has changed
Anthropic restores Claude Fable 5 after refining safety filters, ending an 18-day suspension.

Anthropic has restored access to its Claude Fable 5 model, bringing an end to an 18-day suspension after the US Department of Commerce withdrew the export controls it had imposed on the model on June 12, according to an official post by Anthropic. Reports suggest the issue was traced to a safety filter, which has now been refined to block a prompting technique previously identified by researchers at Amazon.
The technique flagged by Amazon researchers enabled Claude Fable 5 to identify software vulnerabilities and, in one instance, generate code that could be used to exploit a system. According to reports, Anthropic has now trained a new classifier capable of detecting that specific prompting method in more than 99% of cases. Instead of processing such requests with Fable 5, the system reroutes them to the older and less capable Opus 4.8 model. Anthropic also acknowledged that the updated classifier may inadvertently flag some legitimate coding and debugging requests.
How the new classifier works
The new classifier focuses on detecting the reported prompting technique rather than restricting the model's underlying capabilities. Previously, Anthropic had introduced a safeguard that directly limited the model whenever such prompts were detected. That approach drew criticism from researchers, who argued that it unnecessarily reduced the model's capabilities. Under the revised system, suspicious requests are identified and rerouted, rather than disabling Fable 5's functionality altogether.
However, detection-based safeguards were also what researchers managed to bypass when triggering the original restrictions. A classifier trained to detect one known technique offers little protection against new jailbreak methods that may emerge. Anthropic has acknowledged that no frontier AI model can be made completely resistant to jailbreaks and expects additional techniques to surface over time.
Findings from the previous review
Anthropic's assessment, conducted in collaboration with the government and Amazon, found that several leading AI models—including OpenAI's GPT-5.5, Anthropic's Opus 4.8 and Z.ai's Kimi K2.7—were able to identify the same software vulnerabilities. The company said every model it tested, including Haiku 4.5, Sonnet 4.6 and several Opus variants, successfully reproduced the exploit from a single demonstration, reinforcing its conclusion that claims surrounding Mythos-class cyber capabilities had been overstated.
Benchmark performance after its return
Following its restoration, Claude Fable 5 reclaimed benchmark positions that had temporarily been occupied by Z.ai's GLM 5.2 while the model was unavailable. It also achieved the highest publicly accessible score on the Aider Polyglot benchmark's multi-week task evaluation.
To strengthen its security testing, Anthropic has also launched a hacker bounty program, inviting researchers to report new Claude Fable 5 jailbreak techniques. The company added that it will provide designated government partners with early access to future fro

Google burns cash for first time as AI spending pushes 2026 capex to $205 billion
Tesla sells more cars but makes less money: Why Musk’s AI pivot is squeezing profits
OpenAI launches Presence to bring AI agents into customer support and enterprise workflows
Samsung Galaxy Fold 8 Ultra, Fold 8, Flip 8 launched: Here is how much it costs in India with discounts
Florida pastor sues OpenAI, says ChatGPT's medical advice delayed emergency treatment: Report
