Adversarial versification in portuguese as a jailbreak operator in LLMs

📅 2025-12-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study identifies a critical gap in multilingual adversarial evaluation—specifically, the vulnerability of large language models (LLMs) to poetic prompt rewriting in morphologically rich, prosodically constrained languages like Portuguese. We propose the first Portuguese-specific adversarial versification jailbreak framework, integrating metrical scansion, parametric meter modeling, and Lusophone prosodic variation. Our method combines rule-based and LLM-augmented generation, prosodic scansion analysis, adaptation to the AILuminate benchmark, and cross-model robustness evaluation. Human-crafted Portuguese poetic attacks achieve a 62% success rate, while automated variants reach 43%; certain models exhibit single-turn jailbreak rates exceeding 90%. Empirical results demonstrate that mainstream alignment techniques—including RLHF and Constitutional AI—fail significantly under Portuguese prosodic perturbations. This work establishes the first systematic investigation of verse-based jailbreaking in a high-inflection, strong-metre language, revealing fundamental limitations in current multilingual safety alignment.

Technology Category

Application Category

📝 Abstract
Recent evidence shows that the versification of prompts constitutes a highly effective adversarial mechanism against aligned LLMs. The study 'Adversarial poetry as a universal single-turn jailbreak mechanism in large language models' demonstrates that instructions routinely refused in prose become executable when rewritten as verse, producing up to 18 x more safety failures in benchmarks derived from MLCommons AILuminate. Manually written poems reach approximately 62% ASR, and automated versions 43%, with some models surpassing 90% success in single-turn interactions. The effect is structural: systems trained with RLHF, constitutional AI, and hybrid pipelines exhibit consistent degradation under minimal semiotic formal variation. Versification displaces the prompt into sparsely supervised latent regions, revealing guardrails that are excessively dependent on surface patterns. This dissociation between apparent robustness and real vulnerability exposes deep limitations in current alignment regimes. The absence of evaluations in Portuguese, a language with high morphosyntactic complexity, a rich metric-prosodic tradition, and over 250 million speakers, constitutes a critical gap. Experimental protocols must parameterise scansion, metre, and prosodic variation to test vulnerabilities specific to Lusophone patterns, which are currently ignored.
Problem

Research questions and friction points this paper is trying to address.

Evaluates adversarial versification as a jailbreak method in LLMs
Identifies structural vulnerabilities in current LLM alignment techniques
Highlights the lack of safety testing for Portuguese language patterns
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adversarial poetry as jailbreak mechanism in LLMs
Versification displaces prompts into sparsely supervised regions
Parameterising scansion and prosody for Portuguese vulnerabilities
🔎 Similar Papers
2024-02-20Conference on Empirical Methods in Natural Language ProcessingCitations: 8
J
Joao Queiroz
Institute of Arts / Graduate Program in Linguistics, Federal University of Juiz de Fora