On-Policy Distillation for Vision-Language Model Adaptation, an Effective Paradigm on Low-Quality Multimodal Data

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出OnPoKD框架,通过动态调整目标构建策略解决低质量多模态数据下视觉-语言模型适应性问题。
📝 Abstract
Knowledge distillation offers an efficient route to transfer a task-adapted vision-language teacher to a compact student. The training target in current vision-language distillation methods is typically constructed from the teacher prediction and applied uniformly to all training samples, making it unreliable under class and domain shifts. In this paper, we argue that distillation target construction should be treated as a dynamic training decision rather than a fixed recipe. To this end, we propose OnPoKD, an on-policy distillation framework for vision-language model adaptation. To the best of our knowledge, OnPoKD is the first framework that applies on-policy distillation to vision-language model adaptation by learning target construction as a policy decision. OnPoKD learns a lightweight controller that constructs sample-wise adaptive targets using reliability and disagreement cues from the teacher model, student model, and zero-shot prior. Instead of relying on a fixed teacher prediction, the controller dynamically balances teacher supervision, zero-shot prior guidance, and hard-label anchoring through bounded policy actions, allowing the distillation target to adapt to varying sample reliability and training stages. The policy controller is updated with validation feedback, encouraging target construction to optimize transferability rather than merely fitting the training distribution. Since the controller is only used during training, OnPoKD can be seamlessly integrated into existing vision-language distillation pipelines while preserving the original inference architecture and test-time cost. Extensive experiments on Base-to-novel generalization and Cross-dataset transfer benchmarks show that OnPoKD consistently improves over strong vision-language distillation baselines.
Problem

Research questions and friction points this paper is trying to address.

vision-language distillation
dynamic training decision
sample reliability
domain shifts
on-policy distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

on-policy distillation
vision-language model adaptation
sample-wise adaptive targets
policy controller
reliability and disagreement cues
🔎 Similar Papers
No similar papers found.
Hongyuan Zhang
Hongyuan Zhang
The University of Hong Kong
Representation LearningMultimodal LearningGraph Neural NetworksOptimization
Xianda Guo
Xianda Guo
PhD Student at Wuhan University
Stereo Matching, Depth Estimation,Gait Recognition
Y
Yanlun Peng
Great Wall Motor, China
Q
Qianlong Yang
School of Science, China University of Petroleum (East China), China
Y
Yubin Guo
School of Computer Science and Technology, University of Science and Technology of China, China
P
Pinhan Fu
Great Wall Motor, China
Mulin Chen
Mulin Chen
Northwestern Polytechnical University
X
Xiaozhen Qiao
School of Information Science and Technology, University of Science and Technology of China, China
Ping Luo
Ping Luo
National University of Defense Technology
distributed_computing