Institution profile

ServiceNow

Industry researchnorthamerica · us
Official website
Research library131linked papers
Opportunities142open roles
Selected work

Representative Papers

Recent publications

Latest Papers

Backdoor Decontamination Dynamics in LLM Agents

Aug 11, 2026

This study addresses the vulnerability of open-source large language model agents to stealthy backdoors implanted during fine-tuning, which are difficult to detect when trigger conditions remain unobserved. The work systematically investigates the efficacy of defensive poisoning and unlearning in mitigating unknown backdoors, revealing for the first time that trigger recognition and malicious execution can be behaviorally decoupled. It proposes a novel strategy: applying defensive poisoning with analogous triggers followed by depoisoning, which nearly eliminates the original backdoor. Evaluated on the AgentDyn framework with J-lens representation visualization across 115 experiments, defensive poisoning alone removes approximately 56% of backdoors, while combining it with depoisoning achieves near-complete (≈100%) removal. Notably, in multi-backdoor settings, neutralizing one known backdoor incidentally eradicates 87% of coexisting unknown backdoors.

0 citationsRead paper