A famous AI red team expert claimed developing a universal jailbreak that can work against top large language models, like the very secure GPT-5.6 Sol, Claude Opus 5, and Fable.
In a public post on X, Pliny the Liberator described the technique as effective “on ALL models” and across every category he tested. He argued that, because of how the method works, it may be extremely difficult or even impossible to fully patch.
Jailbreak on Top AI Models
Pliny said he will not share the full method right now, unlike other jailbreak releases that go open source right away. He wants a time for responsible sharing so AI labs, security teams, safety researchers, and lawmakers can look at the problem before it gets too common.
He asked AI experts in red teaming, security, alignment, and policy to reach out to him privately. He said this was because of the current political situation and a wish to prevent stricter model rules or bans that might come after a messy public release.
Jailbreaks are prompts or ways of interacting with a model that allow it to ignore its safety rules and produce unsafe results. A universal claim stands out because most bypasses are specific to one model and become stronger after being revealed.
If the technique holds up under independent testing, it would underscore ongoing gaps in:
Safety training and refusal behavior
Guardrail robustness under adversarial prompting
Cross-model generalization of attack patterns
How vendors coordinate fixes without over-blocking legitimate use
Pliny said he does not believe public release would make the world “any more dangerous,” but he acknowledged that others may disagree. During the disclosure period, he aims to map the full impact, measure how much extra capability the method unlocks, and help frame the issue for decision-makers.
Security teams and AI product owners should see this as a warning, not a certain fact. Getting independent checks, vendor advice, and patch instructions will be more important than just the first claim.
Organizations that use these models should keep regular controls until labs respond or the method is properly documented. These controls include checking outputs, giving limited access to tools, having humans review high-risk workflows, and having clear steps for dealing with policy violations.
The researcher said he wants to share the method “when the time is right.” For now, the next step in the industry—private testing or public panic—will affect how this story goes.
Related post
“sockpuppeting” can jailbreak 11 AI models like ChatGPT, Claude, and Gemini
InfoSecBulletin Cybersecurity for mankind
