SpikeOPD: Stable On-Policy Distillation for Autoregressive Spiking Language Models

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出SpikeOPD方法,通过自生成前缀的在线策略蒸馏和一系列稳定措施,解决了脉冲神经网络语言模型训练难的问题。
📝 Abstract
Spiking neural networks (SNNs) offer a path to energy-efficient language modeling through sparse encoding and event-driven computation, but training capable spiking language models from scratch remains difficult. A practical alternative is ANN-to-SNN migration through knowledge distillation (KD), where a pretrained artificial neural network (ANN) teacher supervises an SNN student. Existing migration approaches distill on fixed corpus prefixes, whereas autoregressive inference conditions on self-generated prefixes, creating prefix-source mismatch. It manifests as output-policy mismatch with the ANN teacher and internal spiking-dynamics drift between self-generated and matched corpus prefixes. On-policy distillation (OPD) offers a natural way to mitigate both manifestations by continuing teacher supervision on self-generated prefixes. We evaluate a teacher-only full-KL variant, Vanilla OPD, via a controlled stress test and observe it may suffer from delayed rollout-feedback collapse. This result shows that on-policy coverage alone does not ensure stable adaptation. Motivated by these findings, we propose SpikeOPD, a stable on-policy distillation framework for autoregressive SNNs that learns from self-generated prefixes while maintaining rollout stability. It applies full-KL teacher correction to reduce output-policy mismatch, while matched-prefix policy anchoring constrains policy departure from the frozen reference SNN on the same prefixes. Layerwise spike regularization further limits firing-rate deviations during on-policy adaptation. Across three model scales, SpikeOPD improves average accuracy over the corresponding KD SNNs by 0.8, 1.7, and 2.9 points at 0.125B, 0.35B, and 1.3B, respectively, while preserving their sparse-compute profiles.
Problem

Research questions and friction points this paper is trying to address.

Spiking Neural Networks
Knowledge Distillation
Prefix-Source Mismatch
Autoregressive Inference
On-Policy Distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-policy Distillation
Spiking Neural Networks
Policy Anchoring
Layerwise Spike Regularization
Autoregressive Inference
🔎 Similar Papers
No similar papers found.
E
Enqiao Lu
The Chinese University of Hong Kong, Shenzhen
Xingrui Yu
Xingrui Yu
Scientist, CFAR, A*STAR
Machine LearningRobust Imitation LearningTrustworthy AI
Yiwei Fu
Yiwei Fu
GE Research
Deep LearningReinforcement LearningSelf-Supervised LearningTime Series
Z
Zhenglin Wan
National University of Singapore
P
Pengfei Zhou
National University of Singapore
Wangbo Zhao
Wangbo Zhao
National University of Singapore
Efficient Deep LearningDynamic Neural NetworkMultimodal Model
M
Muqing Jian
The Chinese University of Hong Kong, Shenzhen
X
Xueyi Zhang
National University of Singapore
Yang You
Yang You
Postdoc, Stanford University
3D visioncomputer graphicscomputational geometry
I
Ivor Tsang
Agency for Science, Technology and Research (A*STAR), Singapore