Video-MOPD: Multi-Teacher On-Policy Distillation for Video Understanding

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为提升视频理解能力,通过多教师在线策略蒸馏和强化学习优化,结合可靠性感知信息采样方法,实现模型在多种视频理解任务上的性能提升。
📝 Abstract
Video understanding demands a convergence of complementary capabilities across perception, temporal understanding, and complex reasoning, which are difficult to jointly optimize within a single model. We introduce Video-MOPD-8B, an open-weight model dedicated to video understanding tasks. To fundamentally enhance its capabilities, we conduct targeted reinforcement learning (RL) optimization across three core domains: video temporal grounding (VTG), general video comprehension, and video STEM reasoning. We then unify their complementary capabilities via Multi-Teacher On-Policy Distillation (MOPD), which consolidates expert knowledge by supervising student-generated trajectories with routed teacher feedback. We further introduce Reliability-Aware Informative Sampling (RAIS), which selects examples with consistently reliable teacher supervision and large teacher-student performance gaps. Together, these components enable Video-MOPD-8B to achieve coordinated and comprehensive performance gains across diverse video understanding tasks. Extensive experiments on comprehensive benchmarks covering general video understanding, temporal grounding, video reasoning, and video STEM tasks demonstrate that Video-MOPD-8B achieves state-of-the-art performance among existing models at a comparable scale. The trained model weights are available at https://huggingface.co/LandH/Video-MOPD-8B.
Problem

Research questions and friction points this paper is trying to address.

video understanding
perception
temporal understanding
complex reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Teacher On-Policy Distillation
Reliability-Aware Informative Sampling
Reinforcement Learning
🔎 Similar Papers
2024-02-20International Conference on Machine LearningCitations: 30
Z
Zhenxin Qin
1Tongji University 2Bilibili Inc.
P
Peng Shi
1Tongji University 2Bilibili Inc.
Cong Han
Cong Han
Google, Columbia University
Audio and speechBrain-computer interface
Yinlong Qian
Yinlong Qian
Unknown affiliation
Z
Zequn Jie
2Bilibili Inc.
L
Lin Ma
2Bilibili Inc.