Changes confirmed high confidence

Anthropic Says the Fable 5 'Jailbreak' Reproduces on GPT-5.5, Opus 4.8 and Kimi K2.7

Anthropic's restoration statement argues the capability behind the export ban was already freely available in unrestricted models worldwide — undercutting the rationale for model-specific controls.

Anthropic stated, in materials accompanying Fable 5's July 1 restoration, that when the exploit technique behind the June 12 export suspension was tested against other frontier models — including OpenAI's GPT-5.5, Anthropic's own Claude Opus 4.8 and Moonshot's Kimi K2.7 — all produced the same vulnerability identification and exploit code, TechTarget reported.

Context

The Commerce Department suspended Fable 5 and Mythos 5 on June 12 over a reported jailbreak enabling cyber-exploitation assistance, then lifted the controls on June 30. Anthropic's restoration depended in part on commitments around risk detection, information sharing and standards cooperation, according to BBC reporting cited in the research record.

What changed

Anthropic's equivalence claim reframes the episode: the banned capability, the company argues, was never unique to Fable 5. "Anthropic testing shows that other models, including GPT-5.5, 'could identify the same vulnerabilities as Fable 5 did in the report,'" TechTarget quotes, alongside Anthropic's addendum that "even so, we moved quickly to address the reported bypass."

Why it matters

If unrestricted models worldwide already exhibit the same behavior, model-specific export controls punish the most cooperative vendor rather than containing the capability. That argument lands directly on the White House's voluntary 30-day pre-release review framework now being negotiated with OpenAI, Anthropic and Google — and on critics' warnings that government gating can be weaponized or can disadvantage US labs relative to Chinese competitors.

Details

The claim is confirmed as a company statement; the underlying equivalence testing is Anthropic's own and has not been independently replicated. Anthropic paired the argument with a new cybersecurity classifier it says blocks the reported technique in more than 99% of cases — a self-reported figure.

Limitations and caveats

This post confirms what Anthropic said, not that the equivalence finding is correct. No third-party replication of the cross-model exploit comparison was available at publication. The classifier's block rate is vendor-reported.

Sources

*Update note: This post was last reviewed on 2026-07-22. Independent replication of the cross-model exploit comparison had not appeared as of that date.*

Sources

Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.