Researcher Claims Universal Jailbreak Prompt Simultaneously Bypasses GPT-5.6 and Claude Opus 5

On July 25, 2026, AI red-teaming researcher Pliny the Liberator claimed to have developed a universal jailbreak technique capable of bypassing safety alignment mechanisms in multiple frontier models, including GPT-5.6, Claude Opus 5, and Fable 5.

On July 25, 2026, AI red-teaming researcher Pliny the Liberator claimed to have developed a universal jailbreak technique capable of simultaneously bypassing the safety alignment mechanisms of several frontier models, including GPT-5.6, Claude Opus 5, and Fable 5.

Factual Reconstruction

Pliny the Liberator issued a statement on July 24, 2026, describing the technique as effective against "ALL models" and covering every category he tested. He chose a responsible disclosure path, inviting private contact from experts in AI red-teaming, safety, alignment, and policy, without revealing the specific method.

Mechanism Breakdown

Jailbreak prompts typically bypass a model's safety filters through prompt engineering, causing it to generate prohibited or high-risk outputs. The statement emphasized that due to the method's workings, it may be extremely difficult or impossible to fully patch. Pliny's typical approach includes role-playing and iterative refinement of prompts to gradually erode defenses. A report from the UK AI Safety Institute noted that the institute has identified universal jailbreaks applicable to every system it tested, with these attacks reliably extracting violating information at levels close to those of unguarded models. Pliny stated that this method may not increase world risk levels, but acknowledged that others might hold different views.

Industry Impact

For model providers, this incident reveals continued gaps in the robustness of safety training, refusal behaviors, and guardrails when facing adversarial prompts. OpenAI mentioned at the launch of GPT-5.6 that the model family employs layered protections, real-time checks, and trust- and risk-based access controls, but the statement did not confirm the vulnerability. For developers, relying on frontier model APIs without treating guardrails as part of a comprehensive safety plan poses risks. Any startup that treats refusal behavior as a safety boundary assumes more risk than it acknowledges to its customers. For enterprise users, applications involving customer support, coding, payment processes, or internal data should retain output monitoring, least-privilege tool access, and human review for high-risk workflows.

Parallels and Precedents

The SANS Institute announced Pliny as a keynote speaker for its AI Cybersecurity Summit on April 20–21, 2026, describing him as an anonymous hacker skilled at cracking major AI models shortly after their release and publishing techniques in a GitHub repository that has garnered over 10,000 stars. SANS also noted that Pliny leads the BT6 White Hat Collective, and TIME included him in its 2025 TIME100 AI list. A report from the UK AI Safety Institute shows that universal jailbreaks have emerged in every system it tested, echoing Pliny's statement.

Strategic Assessment

During private review, labs will evaluate the method's impact scope and decide whether to coordinate patches or adjust disclosure processes. Enterprise users should continue maintaining existing controls rather than relying on a single guardrail. The above assessment is based on current statements and historical patterns, constituting analysis rather than fact.