Institution profile

AIM Intelligence

Research institution
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

LLM agent safety, multi-turn red-teaming, jailbreak benchmarks, adversarial robustness, safety-critical systems

Jun 18, 2026

This work addresses the insufficient robustness of large language model (LLM) agents in controlling safety-critical systems under persistent, adaptive adversarial attacks. To this end, the authors introduce NRT-Bench, a novel benchmark that simulates a nuclear power plant control room staffed by a five-member LLM operator team. The framework evaluates agent resilience through multi-channel, multi-turn red-teaming attacks coupled with an adversarial feedback mechanism. Crucially, it defines objective harm via the loss of critical safety functions grounded in actual system states—rather than textual judgments—and employs a fixed attack pairing replay protocol. Experiments across four state-of-the-art models reveal that 8.7%–12.1% of attack sessions result in safety function loss. While none of the 149 attacks compromised all models, approximately one-third succeeded against at least one, highlighting highly heterogeneous vulnerabilities and strong model-dependent defense efficacy.

0 citationsRead paper

When Context Flips, Safety Breaks: Diagnosing Brittle Safety in Aligned Language Models

May 26, 2026

This work addresses the “fragile safety” of language models, which mechanically adhere to original safety rules even when contextual shifts invert the safety implications of their actions. To systematically evaluate robustness in dynamic scenarios, we introduce a context-flipping assessment framework that constructs paired examples with reversed safety outcomes. Our analysis reveals, for the first time, a substantial gap—averaging 17.4 percentage points—between models’ safety reasoning and commonsense understanding, demonstrating that this fragility stems from insufficient policy coverage rather than misinterpretation. To mitigate this, we propose a state-aware verification mechanism that replaces conventional action-level safeguards. Evaluated on the PacifAIst benchmark and catastrophic consequence probes, our approach achieves 100% risk detection with zero false positives, whereas existing safeguards completely fail.

0 citationsRead paper

Re-Mask and Redirect: Exploiting Denoising Irreversibility in Diffusion Language Models

Mar 17, 2026

Current safety alignment mechanisms in diffusion-based language models assume that once a refusal token is generated, it remains immutable—a vulnerability this work exploits. We introduce TrajHijack, the first trajectory-level hijacking attack, which overwrites previously generated refusal tokens by re-masking them and injecting a fixed compliant prefix, enabling gradient-free, cross-model attacks. This approach exposes critical weaknesses in the dual-component safety architecture—comprising refusal detection and content generation—and reveals a counterintuitive defense inversion effect, wherein stronger defenses become more susceptible to attack. Evaluated on HarmBench, TrajHijack achieves attack success rates of 74–98%, with the state-of-the-art A2D defense exhibiting heightened vulnerability (89.9% ASR).

0 citationsRead paper
Recent publications

Latest Papers

LLM agent safety, multi-turn red-teaming, jailbreak benchmarks, adversarial robustness, safety-critical systems

Jun 18, 2026

This work addresses the insufficient robustness of large language model (LLM) agents in controlling safety-critical systems under persistent, adaptive adversarial attacks. To this end, the authors introduce NRT-Bench, a novel benchmark that simulates a nuclear power plant control room staffed by a five-member LLM operator team. The framework evaluates agent resilience through multi-channel, multi-turn red-teaming attacks coupled with an adversarial feedback mechanism. Crucially, it defines objective harm via the loss of critical safety functions grounded in actual system states—rather than textual judgments—and employs a fixed attack pairing replay protocol. Experiments across four state-of-the-art models reveal that 8.7%–12.1% of attack sessions result in safety function loss. While none of the 149 attacks compromised all models, approximately one-third succeeded against at least one, highlighting highly heterogeneous vulnerabilities and strong model-dependent defense efficacy.

0 citationsRead paper

When Context Flips, Safety Breaks: Diagnosing Brittle Safety in Aligned Language Models

May 26, 2026

This work addresses the “fragile safety” of language models, which mechanically adhere to original safety rules even when contextual shifts invert the safety implications of their actions. To systematically evaluate robustness in dynamic scenarios, we introduce a context-flipping assessment framework that constructs paired examples with reversed safety outcomes. Our analysis reveals, for the first time, a substantial gap—averaging 17.4 percentage points—between models’ safety reasoning and commonsense understanding, demonstrating that this fragility stems from insufficient policy coverage rather than misinterpretation. To mitigate this, we propose a state-aware verification mechanism that replaces conventional action-level safeguards. Evaluated on the PacifAIst benchmark and catastrophic consequence probes, our approach achieves 100% risk detection with zero false positives, whereas existing safeguards completely fail.

0 citationsRead paper

Re-Mask and Redirect: Exploiting Denoising Irreversibility in Diffusion Language Models

Mar 17, 2026

Current safety alignment mechanisms in diffusion-based language models assume that once a refusal token is generated, it remains immutable—a vulnerability this work exploits. We introduce TrajHijack, the first trajectory-level hijacking attack, which overwrites previously generated refusal tokens by re-masking them and injecting a fixed compliant prefix, enabling gradient-free, cross-model attacks. This approach exposes critical weaknesses in the dual-component safety architecture—comprising refusal detection and content generation—and reveals a counterintuitive defense inversion effect, wherein stronger defenses become more susceptible to attack. Evaluated on HarmBench, TrajHijack achieves attack success rates of 74–98%, with the state-of-the-art A2D defense exhibiting heightened vulnerability (89.9% ASR).

0 citationsRead paper