Accelerating Visual On-Policy Distillation with Batched Speculative Jacobi Rollouts

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决视觉在线策略蒸馏中自回归解码成本高的问题,提出HB-SJD方法,通过批处理投机雅可比展开来加速训练过程。
📝 Abstract
Visual on-policy distillation (OPD) improves the training of compact visual autoregressive models by learning from trajectories generated by the current student. However, these online rollouts are still produced token by token with autoregressive decoding, which adds substantial cost to every on-policy training step. Speculative Jacobi Decoding (SJD) provides an alternative because it can process multiple tokens in parallel without an auxiliary draft model, but the original method is designed for single-sequence inference. We introduce HB-SJD, a batched SJD rollout backend for visual OPD. HB-SJD allows each image to advance independently according to its own decoding progress, while images at different sequence positions are still verified in batched model forwards. As images finish, HB-SJD switches between Full and Compact execution to reduce the cost of later rollout rounds. HB-SJD only replaces the student rollout backend and leaves the teacher, distillation objective, and optimization procedure unchanged. Experiments with LlamaGen show that HB-SJD substantially reduces rollout and end-to-end training time while preserving the generation quality of the distilled student.
Problem

Research questions and friction points this paper is trying to address.

Visual On-Policy Distillation
Autoregressive Decoding
Rollout Cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

Batched Speculative Jacobi Rollouts
Visual On-Policy Distillation
Parallel Token Processing
🔎 Similar Papers
2024-07-21arXiv.orgCitations: 1
B
Bingqi Shan
Shenzhen Key Laboratory of Internet Information Collaboration, Harbin Institute of Technology, Shenzhen, Shenzhen, China
Z
Zhehao Yu
Shenzhen Key Laboratory of Internet Information Collaboration, Harbin Institute of Technology, Shenzhen, Shenzhen, China
K
Kenhong Lin
Shenzhen Key Laboratory of Internet Information Collaboration, Harbin Institute of Technology, Shenzhen, Shenzhen, China
Baoquan Zhang
Baoquan Zhang
Harbin Institute of Technology, Shenzhen
knowledge-guided machine learningmeta-learningfew-shot learning