Training Trajectories Determine Circuit Removability in Annealable Soft-Prior Transformers

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过可退火软先验Transformer测试了在移除先验后电路是否仍能正常工作,发现训练轨迹对小规模离散检索任务中的电路可移除性有重要影响。
📝 Abstract
Soft positional priors can help small Transformers learn retrieval circuits, but it is unclear whether the resulting circuits remain functional once the prior is removed. We test this with an annealable soft-prior Transformer whose attention biases can be learned, faded, or zeroed during training and evaluation. On associative recall, unforced models perform well with the prior active ($0.772 \pm 0.020$) but collapse at zero gate ($0.095 \pm 0.009$). Smooth fade-to-zero training preserves high zero-gate accuracy ($0.734 \pm 0.028$), whereas forced-zero training, hard switching, and post hoc continuation fail to recover the same effect. The pattern also appears on Markov induction. Linear regression ICL provides a boundary case because zero-gate training can learn that task directly. Mechanistic traces show that circuit consolidation occurs after the gate reaches zero, even though the responsible heads vary across seeds. These results suggest that circuit removability in small discrete retrieval tasks depends on the training trajectory, not just the final architecture.
Problem

Research questions and friction points this paper is trying to address.

soft positional priors
transformers
circuit removability
training trajectory
Innovation

Methods, ideas, or system contributions that make the work stand out.

annealable soft-prior Transformer
training trajectory
circuit removability
smooth fade-to-zero training
🔎 Similar Papers