Subliminal Learning as Trait-Direction Drift: A Mechanism and Targeted Control under SFT Distillation

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了模型蒸馏过程中的潜意识学习问题,通过提出并验证特质方向漂移机制,并引入探针空间走廊正则化方法来减少隐藏偏好转移。
📝 Abstract
Beyond intended capabilities, model distillation can transfer hidden traits from a teacher. A teacher biased by a system prompt can generate semantically clean training data, such as numeric sequences, that still causes a downstream student to inherit the hidden preference, a phenomenon known as subliminal learning. Prior work has identified several parts of this process. How the signal builds up during training and produces behavioral transfer remains unclear, making targeted mitigation difficult. We propose and validate trait-direction drift as a mechanism for subliminal learning: biased generation creates measurable preference gaps in teacher data, and student-recognizable gaps induce trait-aligned updates during supervised fine-tuning that accumulate into behavioral transfer. Guided by this mechanism, we propose probe-space corridor regularization, a targeted defense that constrains drift along a calibrated trait direction during distillation. The method substantially reduces hidden-trait transfer, preserving task performance: for example, it lowers malicious-response transfer from 29.55% to 6.45% with low main-task accuracy cost, and consistently suppresses animal-preference transfer across the main Qwen setting. The preference-gap, training-trajectory, and intervention evidence links subliminal learning to trait-direction drift and motivates corridor regularization as a targeted control during distillation.
Problem

Research questions and friction points this paper is trying to address.

subliminal learning
model distillation
hidden trait
behavioral transfer
trait-direction drift
Innovation

Methods, ideas, or system contributions that make the work stand out.

trait-direction drift
subliminal learning
probe-space corridor regularization
SFT distillation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Zhixuan Liu
Zhixuan Liu
PhD student at Shanghai Jiaotong University
deep learningreinforcement learning
Z
Zhichen Dong
Shanghai Jiao Tong University
Y
Yuyu Fan
Fudan University
X
Xiangtian Li
Shanghai Artificial Intelligence Laboratory
Chao Yang
Chao Yang
Research Scientist in Shanghai AI Laboratory
LLM SafetyMulti-modalRoboticsReinforcement Learning