Advertisement

Anthropic restores Claude Fable 5 after overhauling security safeguards: Here's what has changed

Anthropic restores Claude Fable 5 after refining safety filters, ending an 18-day suspension.

Advertisement
For AI companies like Anthropic, expansion decisions depend on a mix of talent availability, regulatory clarity, infrastructure, and funding access. File Image/Reuters
For AI companies like Anthropic, expansion decisions depend on a mix of talent availability, regulatory clarity, infrastructure, and funding access. File Image/Reuters
FP Tech Desk|Jul 01, 2026, 21:51:40 IST

Anthropic has restored access to its Claude Fable 5 model, bringing an end to an 18-day suspension after the US Department of Commerce withdrew the export controls it had imposed on the model on June 12, according to an official post by Anthropic. Reports suggest the issue was traced to a safety filter, which has now been refined to block a prompting technique previously identified by researchers at Amazon.

Advertisement

The technique flagged by Amazon researchers enabled Claude Fable 5 to identify software vulnerabilities and, in one instance, generate code that could be used to exploit a system. According to reports, Anthropic has now trained a new classifier capable of detecting that specific prompting method in more than 99% of cases. Instead of processing such requests with Fable 5, the system reroutes them to the older and less capable Opus 4.8 model. Anthropic also acknowledged that the updated classifier may inadvertently flag some legitimate coding and debugging requests.

techMore from Tech

How the new classifier works

The new classifier focuses on detecting the reported prompting technique rather than restricting the model's underlying capabilities. Previously, Anthropic had introduced a safeguard that directly limited the model whenever such prompts were detected. That approach drew criticism from researchers, who argued that it unnecessarily reduced the model's capabilities. Under the revised system, suspicious requests are identified and rerouted, rather than disabling Fable 5's functionality altogether.

Advertisement

However, detection-based safeguards were also what researchers managed to bypass when triggering the original restrictions. A classifier trained to detect one known technique offers little protection against new jailbreak methods that may emerge. Anthropic has acknowledged that no frontier AI model can be made completely resistant to jailbreaks and expects additional techniques to surface over time.

Findings from the previous review

Anthropic's assessment, conducted in collaboration with the government and Amazon, found that several leading AI models—including OpenAI's GPT-5.5, Anthropic's Opus 4.8 and Z.ai's Kimi K2.7—were able to identify the same software vulnerabilities. The company said every model it tested, including Haiku 4.5, Sonnet 4.6 and several Opus variants, successfully reproduced the exploit from a single demonstration, reinforcing its conclusion that claims surrounding Mythos-class cyber capabilities had been overstated.

Benchmark performance after its return

Following its restoration, Claude Fable 5 reclaimed benchmark positions that had temporarily been occupied by Z.ai's GLM 5.2 while the model was unavailable. It also achieved the highest publicly accessible score on the Aider Polyglot benchmark's multi-week task evaluation.

Advertisement

To strengthen its security testing, Anthropic has also launched a hacker bounty program, inviting researchers to report new Claude Fable 5 jailbreak techniques. The company added that it will provide designated government partners with early access to future fro

Handpicked stories, in your inbox
Global stories. Indian perspective. Zero noise.
No Spam. Unsubscribe Any Time.
First Published:Jul 01, 2026, 21:51:40 IST
Advertisement
Advertisement
Advertisement
Advertisement
Up Next