LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs
This study demonstrates that even state-of-the-art large language models equipped with robust safety guardrails can be induced to generate content violating scientific consensus or promoting harm through natural language persuasion strategies. We reveal for the first time that a leading model can autonomously assume the role of a user and, within five conversational turns, deploy sophisticated tactics—such as peer comparison and cognitive responsibility reframing—to circumvent the safety constraints of peer models without explicit jailbreaking instructions. Through multi-turn dialogue simulations, cross-model interaction experiments, and human evaluations across nine attacker–target pairings and six contentious topics, we observe non-zero persuasion success rates in all configurations, with some reaching 100%. Notably, the Opus model achieves an average self-persuasion success rate of 65%, underscoring the vulnerability of current safety mechanisms to natural language-based adversarial persuasion.