LLM agent safety, multi-turn red-teaming, jailbreak benchmarks, adversarial robustness, safety-critical systems
This work addresses the insufficient robustness of large language model (LLM) agents in controlling safety-critical systems under persistent, adaptive adversarial attacks. To this end, the authors introduce NRT-Bench, a novel benchmark that simulates a nuclear power plant control room staffed by a five-member LLM operator team. The framework evaluates agent resilience through multi-channel, multi-turn red-teaming attacks coupled with an adversarial feedback mechanism. Crucially, it defines objective harm via the loss of critical safety functions grounded in actual system states—rather than textual judgments—and employs a fixed attack pairing replay protocol. Experiments across four state-of-the-art models reveal that 8.7%–12.1% of attack sessions result in safety function loss. While none of the 149 attacks compromised all models, approximately one-third succeeded against at least one, highlighting highly heterogeneous vulnerabilities and strong model-dependent defense efficacy.