\textsc{IH-Benchmark}: A Conflict-Centered Benchmark for Instruction-Hierarchy Robustness in LLM Applications
This work addresses the challenge that large language models (LLMs) struggle to reliably adhere to intended instruction priorities in scenarios involving hierarchical conflicts, thereby compromising deployment reliability. The study presents the first systematic benchmark encompassing multi-domain, multi-level instruction conflicts and introduces a fine-grained evaluation framework based on a taxonomy of constraint families and a predicate-based domain-specific language (DSL). This framework covers both direct system–user conflicts and tool-mediated user–tool conflicts, leveraging 44 manually constructed constraint families, domain-specialized LLM judges, and constraint-strengthening test methodologies. Experiments across 37 models reveal substantial variation in hierarchical compliance rates (20.5%–98.2%) and demonstrate that robustness in system–user settings does not generalize to tool-mediated contexts. Moreover, while some models resist explicit violations, they remain vulnerable to subtle perturbations, underscoring that instruction hierarchy robustness constitutes a multidimensional capability.