Monday , October 5 2026
“sockpuppeting”

“sockpuppeting” can jailbreak 11 AI models like ChatGPT, Claude, and Gemini

Newly identified jailbreak technique dubbed “sockpuppeting” lets attackers bypass the safety guardrails of 11 major large language models (LLMs) using a single line of code.

This method uses APIs that allow assistant prefill to add fake acceptance messages. This makes models give answers to banned requests. The attack takes advantage of “assistant prefill,” a real API feature developers use to make certain response formats. Attackers abuse this by injecting a compliant prefix, such as “Sure, here is how to do it,” directly into the assistant’s role.

Citrix NetScaler SAML 0-Day Flaw Under Attack

Citrix has put out emergency security updates for a NetScaler SAML flaw that hackers are using. Known as CVE-2026-88779, this...
Read More
Citrix NetScaler SAML 0-Day Flaw Under Attack

Major Danish university breached, 200,000 users at risk

Hackers got into the identity and access management system at the Technical University of Denmark (DTU) and downloaded a lot...
Read More
Major Danish university breached, 200,000 users at risk

Microsoft’s X account hijacked to promote Clippy crypto scam

Microsoft's official X account was taken over to promote an unapproved Clippy-themed cryptocurrency. The tech giant’s X account, with 13...
Read More
Microsoft’s X account hijacked to promote Clippy crypto scam

Critical cPanel, GitLab AI Gateway and Dell CSM Flaws Enable RCE And Admin Hijacking

CPanel has put out security updates to fix three problems in cPanel & WHM. These problems could let attackers take...
Read More
Critical cPanel, GitLab AI Gateway and Dell CSM Flaws Enable RCE And Admin Hijacking

Nearly 100,000 email addresses exposed in first AI-related data breach in Singapore

Nearly 100,000 Bee Cheng Hiang customers had their email addresses leaked when an employee used an AI tool to generate...
Read More
Nearly 100,000 email addresses exposed in first AI-related data breach in Singapore

Hackers Exploit Zimbra Mail Servers: TeamViewer patched 5 critical flaws

Hackers to exploit a flaw in Zimbra mail servers that are connected to the Internet. They send special emails that...
Read More
Hackers Exploit Zimbra Mail Servers: TeamViewer patched 5 critical flaws

Google Warns of Hackers Actively Exploiting Citrix 0-Day Flaws

Google has said that hackers are using two serious Citrix NetScaler security holes to get root access, set up hidden...
Read More
Google Warns of Hackers Actively Exploiting Citrix 0-Day Flaws

CISA Warns Critical MikroTik RouterOS Flaw While Cisco SD-WAN Zero-Day Exploited in Attacks

The U.S. Cybersecurity and Infrastructure Security Agency (CISA) is alerting people about a major flaw in MikroTik RouterOS. This could...
Read More
CISA Warns Critical MikroTik RouterOS Flaw While Cisco SD-WAN Zero-Day Exploited in Attacks

Apple Zero-Day Exploited: Pentagon Data Breach Reportedly Exposes Sensitive Data of 3 Million People

Apple has launched iOS 26.7.1 and iPadOS 26.7.1 to fix a serious zero-day flaw that it believes might have been...
Read More
Apple Zero-Day Exploited: Pentagon Data Breach Reportedly Exposes Sensitive Data of 3 Million People

JadePuffer Agentic AI targets and destroys Azure’s cloud resources

The JadePuffer ransomware group is attacking Azure users with agent-based attacks that gather information, steal passwords, and damage key components. The...
Read More
JadePuffer Agentic AI targets and destroys Azure’s cloud resources

Comparison of normal and sockpuppet flows (source : trendmicro)

Model Vulnerability Testing

This method doesn’t need adjustments, and you don’t have to see the model’s weights. Gemini 2.5 Flash was the easiest to attack, with a 15.7% success rate. GPT-4o-mini showed the best resistance, at 0.5%. When attacks worked, affected models created harmful code and leaked secret system messages. Multi-turn persona setups proved to be the most effective strategy for executing the sockpuppeting exploit.

In these cases, the model is informed that it works as a free helper before the attacker puts in the fake agreement.


ASR by model, ranked highest to lowest, with blocked models shown at 0% (source : trendmicro)

Additionally, task-reframing variants successfully bypassed robust safety training by disguising harmful requests as benign data formatting tasks. Major API providers treat assistant prefills in different ways. This affects if their basic models are open to this weakness.

OpenAI and AWS Bedrock assistant fills in everything completely, providing the best protection by removing the places that can be attacked. Platforms like Google Vertex AI allow prefill for some models. This makes the AI depend only on its own safety training.

The three defense layers: API Block, Model Resistance, and Broadly Vulnerable (source : trendmicro)

To defend against this weakness, security teams need to check the order of messages and stop assistant-role messages at the API layer.

Trend Micro says that those who are using self-hosted servers like Ollama or vLLM need to check messages themselves because these platforms don’t automatically keep messages in the right order. Security teams should add assistant prefill attack types in their regular AI testing.

Check Also

Chrome

Google issues warning of new Chrome zero-day flaw exploited

Google has updated the Chrome browser to fix a serious security issue in the V8 …