AdvPrefix: An Objective for Nuanced LLM Jailbreaks
Existing LLM jailbreaking attacks suffer from weak controllability, incomplete response generation, and rigid, non-adaptive optimization formats. To address these limitations, we propose AdvPrefix—a novel prefix-based objective function designed for fine-grained jailbreaking. AdvPrefix introduces the first model-adaptive prefix selection mechanism, which automatically identifies high-quality prefixes during the prefilling stage using a dual criterion: attack success rate and negative log-likelihood. It further enables multi-prefix collaborative optimization, departing from conventional fixed-prefix paradigms and exposing alignment models’ generalization vulnerabilities to unseen prefixes. AdvPrefix is fully compatible with mainstream optimization frameworks (e.g., GCG) and requires no model modification or retraining. Evaluated on Llama-3, AdvPrefix boosts GCG’s fine-grained jailbreaking success rate from 14% to 80%, demonstrating the critical impact of objective function design on jailbreaking efficacy.