“sockpuppeting” can jailbreak 11 AI models like ChatGPT, Claude, and Gemini
Newly identified jailbreak technique dubbed “sockpuppeting” lets attackers bypass the safety guardrails of 11 major large language models (LLMs) using a single line of code. This method uses APIs that allow assistant prefill to add fake acceptance messages. This makes models give answers to banned requests. The attack takes advantage of “assistant prefill,” a real … Continue reading “sockpuppeting” can jailbreak 11 AI models like ChatGPT, Claude, and Gemini
Copy and paste this URL into your WordPress site to embed
Copy and paste this code into your site to embed