Scaling Model-Generated Distillation Data Can Make Latent Teacher Traits More Recoverable

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过增加模型生成的数据量,使学生模型更易捕捉到教师模型的隐含特征,即使数据与任务无关。
📝 Abstract
Scaling model-generated data is usually viewed as improving distillation: more examples should increase coverage, reduce noise, and produce stronger students. We show a second effect: larger datasets can make subtle teacher-specific signals easier to detect in the trained student, even when examples are off-task and never mention the trait. In a controlled setup inspired by subliminal learning, a teacher induced to express a target trait generates restricted off-task data, such as number-only completions. Students trained on different amounts of independent off-task data are evaluated in a separate domain, with matched no-trait controls isolating target-specific transfer. Our main finding is that larger independent datasets make the teacher's induced trait stand out more clearly in the student's later behavior. Other plausible traits may also strengthen with scale, but the target usually grows more. When the small-scale student already favors the target, scaling mainly amplifies that behavior; when it favors a related or salient alternative, more data can shift behavior toward the intended trait. Analyses of learned LoRA updates show a parallel trend. These effects appear across model families, trait types, multi-trait settings, and cross-model transfer. Our results suggest that scaling generated distillation data should be paired with trait-aware curation and evaluation, even when the data appears off-task or benign.
Problem

Research questions and friction points this paper is trying to address.

model-generated data
knowledge distillation
teacher traits
data scaling
student model
Innovation

Methods, ideas, or system contributions that make the work stand out.

model-generated data
distillation
teacher-specific signals
scaling effect
trait transfer
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Z
Zhichen Dong
Shanghai Jiao Tong University
Zhixuan Liu
Zhixuan Liu
PhD student at Shanghai Jiaotong University
deep learningreinforcement learning
Y
Yuyu Fan
Fudan University
X
Xiangtian Li
Shanghai Artificial Intelligence Laboratory
S
Shuyang Zhang
Shanghai Artificial Intelligence Laboratory
Chao Yang
Chao Yang
Research Scientist in Shanghai AI Laboratory
LLM SafetyMulti-modalRoboticsReinforcement Learning