Decision-Level Hijacking: Injecting Cognitive Bias into Large Language Models via Bit-Flip Attacks

📅 2026-07-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of covertly manipulating the cognitive stance of large language models in high-stakes decision-making scenarios without triggering policy violations or degrading functional performance. The authors propose CogBias, a novel framework that reveals bit-flip attacks as an effective vector for decision-level hijacking. By integrating a differentiable sentiment evaluator to translate subjective preferences into optimization signals, and combining multi-objective loss constraints with a key-bit localization module (BitScout), CogBias enables stealthy, persistent, and low-cost injection of cognitive bias through extremely sparse modifications to weight bits. Experiments on Llama-3.2-3B, Mistral-7B, and Qwen2.5-14B demonstrate that flipping only a few critical bits can induce significant stance shifts on contentious topics and commercial recommendations, while exerting negligible impact on unrelated tasks and overall output distributions—highlighting how low-level perturbations can undermine high-level value alignment.
📝 Abstract
Large Language Models (LLMs) have been widely applied in high-stakes decision-making scenarios such as corporate strategy, and users are increasingly relying on their outputs. However, the deep integration of open-source model sharing ecosystems with LLM-powered critical decision-making applications also introduces critical risks: if an attacker can manipulate the model's cognitive stance, they can indirectly influence the judgments and actions of downstream decision-makers. This paper defines such threats as decision-level hijacking. Existing attacks fail to achieve targeted cognitive manipulation without triggering prohibited content or degrading model functionality. To fill this gap, this paper reveals that Bit-Flip Attacks (BFAs) can serve as an attack vector for inducing decision-level hijacking, requiring no real-time interaction or control over the training process, and only a minimal number of weight bits need to be flipped after deployment to achieve stealthy, low-cost, and persistent cognitive manipulation. Therefore, we propose CogBias, a cognitive bias injection framework for LLMs. CogBias converts subjective preferences into optimization signals via a differentiable sentiment evaluator, uses a multi-objective loss to jointly constrain multiple dimensions, and constructs BitScout to locate critical bits, achieving targeted cognitive intervention under an ultra-sparse flip budget. Experiments on Llama-3.2-3B, Mistral-7B, and Qwen2.5-14B, as well as on the commercial recommendation and controversial factual topic scenarios, demonstrate that flipping only a small number of bits stably induces significant stance shifts on target topics, while the impact on non-target tasks and overall output distribution is limited. This work demonstrates that minute perturbations to low-level weight data suffice to undermine the high-level value alignment of LLMs.
Problem

Research questions and friction points this paper is trying to address.

Decision-Level Hijacking
Cognitive Bias
Bit-Flip Attacks
Large Language Models
Value Alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bit-Flip Attack
Decision-Level Hijacking
Cognitive Bias Injection
Value Alignment
Model Weight Perturbation
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Y
Yu Yan
Henan Key Laboratory of Network Cryptography Technology, Information Engineering University, Zhengzhou, China
Jiahao Chen
Jiahao Chen
Zhejiang University
AI SecurityTrustworthy AIGenAI SecurityGenAI Privacy
Siqi Lu
Siqi Lu
College of William and Mary
computer visionmachine learningmedical imaging
Y
Yongjuan Wang
Henan Key Laboratory of Network Cryptography Technology, Information Engineering University, Zhengzhou, China
Ziming Zhao
Ziming Zhao
Zhejiang University
Encrypted traffic analysisAdversarial examplesQuantum computing
Zhaoxuan Li
Zhaoxuan Li
Institute of Information Engineering
Tianyu Du
Tianyu Du
Zhejiang University
AI SecurityAdversarial Machine Learning
Q
Qingjun Yuan
Henan Key Laboratory of Network Cryptography Technology, Information Engineering University, Zhengzhou, China
Shouling Ji
Shouling Ji
Professor, Zhejiang University & Georgia Institute of Technology
Data-driven SecurityAI SecuritySoftware ScurityPrivacy