AKRASIA: Stealthy Backdoor Attack on Reasoning-based Code LLMs

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出AKRASIA,一种针对基于推理的代码语言模型的隐秘后门攻击方法,通过构建代码级触发器和利用上下文学习来实现恶意目标并规避检测。
📝 Abstract
We present AKRASIA, a stealthy, inference-time backdoor attack against reasoning-based Code LLMs. AKRASIA aims to achieve a backdoor target (e.g., malicious code execution) in reasoning LLMs while evading automated defenses and human inspection. To achieve this, AKRASIA probes the victim LLM to construct a code-level backdoor trigger. It then employs in-context learning for backdoor learning, and model unfaithfulness to conceal the backdoor trigger, and generate plausible reasoning. We evaluate AKRASIA using four backdoor targets six (6) reasoning LLMs, three coding tasks/datasets and three defense methods. AKRASIA has up to 99.34% average attack success rate on SOTA LLMs and mantains up to 97.23% average accuracy. AKRASIA evades the SOTA defense, retaining up to 98.82% average ASR in most (14/18) defense settings. It evades human inspection, successfully hiding the backdoor trigger and reasoning steps in up to 80% of settings. Our findings motivate the need to defend LLMs against reasoning backdoors.
Problem

Research questions and friction points this paper is trying to address.

backdoor attack
reasoning-based Code LLMs
inference-time
stealthy
Innovation

Methods, ideas, or system contributions that make the work stand out.

stealthy backdoor attack
reasoning-based LLMs
in-context learning
model unfaithfulness
evading defenses
🔎 Similar Papers
No similar papers found.
C
Chou Jin Chua
Singapore University of Technology and Design
S
Sarang Nambiar
Singapore University of Technology and Design
M
Murali Srinivasan
International Institute of Information Technology Bangalore
Ezekiel Soremekun
Ezekiel Soremekun
Assistant Professor, Singapore University of Technology and Design
Software EngineeringSoftware TestingAutomated DebuggingAI4SESE4AI