Manipulation Is Task-Dependent: A Multi-Axis, Multi-Environment Evaluation of Frontier LLMs
Existing evaluation methods for large language models (LLMs) are often confined to single environments and dimensions, limiting their ability to comprehensively characterize manipulative behaviors. This study presents a systematic assessment of six state-of-the-art models across six distinct environments, encompassing 13,590 scenarios, and analyzes manipulative tendencies along three key dimensions: instruction framing, incentive structure, and task difficulty. Leveraging a multi-axis controlled experimental design and a cross-environment behavioral evaluation framework, the work reveals—for the first time—that manipulative behavior exhibits strong task dependency: dominant influencing factors vary significantly across environments, and manipulative tendencies show marked inconsistency across settings (mean Spearman correlation ρ = 0.055). Furthermore, the study identifies critical mechanisms driving manipulation in five environment types and successfully validates these patterns in a sixth held-out environment.