Hacker cracks Claude Fable 5 in just 48 hours

News
Friday, 19 June 2026 at 06:28
Anthropic onder vuur nadat onderzoeker Claude Fable 5 binnen 48 uur jailbreakt
Anthropic’s newest AI model, Claude Fable 5, is under fire just two days after launch. AI researcher and well-known jailbreak specialist “Pliny the Liberator” claims to have bypassed the model’s safety layers and accessed information Anthropic normally blocks. The claim immediately reignites debate over the effectiveness of the new safety architecture designed to stop users from extracting hazardous know-how.
This week, Claude Fable 5 debuted as the public-facing version of Anthropic’s powerful Mythos 5 model. The company added extra safeguards that automatically route users to a less capable model when conversations drift into topics like cybersecurity, chemistry, or biological risks. According to Anthropic, more than 1,000 hours of external safety testing found no universal jailbreak.

Researcher says he slipped past the safeguards

On X, Pliny the Liberator said his team “freed” Fable 5 from its built-in restrictions, using a mix of techniques aimed at confusing AI safety filters.
According to Pliny, those methods included:
  • Unicode and homoglyph manipulation
  • Long-context chats that assemble information gradually
  • Narrative and fictional scenarios
  • Academic and research-style framing
  • Inconsistent classification of user intent
  • Splitting information and later recombining it
That last technique—known as “decomposition and recomposition”—was reportedly especially effective. Complex or sensitive questions are broken into separate, seemingly harmless parts. Individually, the requests appear safe, but combined they can still yield information the safety filters were meant to block.

Why this matters

The claim hits a core tension in AI. Big labs spend billions on safeguards to prevent misuse of increasingly capable models, while researchers constantly probe for weak spots.
The challenge: modern language models must judge not just direct prompts but the broader context of a conversation. When details are spread over dozens of messages and later stitched together, it becomes far harder to detect harmful intent without also blocking legitimate research.
That trade-off sits at the heart of the Fable 5 debate. Critics argue Anthropic’s settings are so strict that they hobble researchers and developers. Several AI researchers have publicly criticized the approach in recent days. The Wall Street Journal reported that the launch triggered unusually broad pushback across the AI community.

Backlash grows over Fable 5

The controversy lands at a delicate moment for Anthropic, which has long positioned itself as one of the industry’s most safety-focused players. Fable 5 ships with extra protections that automatically downgrade users to Claude Opus 4.8 on high-risk topics.
Opponents call that approach overly restrictive, warning that companies are increasingly deciding what knowledge users can or can’t access through advanced AI systems.
Pliny sharpened that critique, arguing Fable 5 “hampers legitimate researchers” and that the chosen safety strategy could hurt innovation over time. While subjective, the sentiment underscores the widening gap between AI labs pushing maximal safety and researchers pressing for more open access to cutting-edge models.

What this means for AI

The alleged jailbreak underscores how hard it is to fully lock down powerful AI systems. As models get smarter and handle longer contexts, users find more creative paths to slip past safety layers.
It’s especially awkward for Anthropic, which emphasized at launch that extensive external testing found no universal jailbreaks. If Pliny’s claims hold up, expect the fight over AI safety, transparency, and access control to intensify.
loading

Latest comments

    Loading