Efficient Reasoning Distillation: Small Video-Language Models via Synthetic CoT and Difficulty-Aware Fine-Tuning

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种通过合成思维链和难度感知微调的方法,将推理能力高效地蒸馏到小型视频-语言模型中,以解决视频问答问题。
📝 Abstract
We present an efficient method to distill reasoning capabilities into compact video-language models (VLMs) for video question answering (VideoQA). Our approach fine-tunes a 2B-parameter model using only $\sim$900 uncertainty-selected examples, each augmented with synthetic chain-of-thought (CoT) rationales generated by a 4B teacher. Despite its minimal compute cost - under two hours on a single A100 GPU - our method enables the 2B model to outperform VLMs up to 4$\times$ larger, and generalize across CinePile, ActivityNet-QA, and MLVU, approaching the performance of its own 4B teacher. A key finding is that placing CoT rationales after the answer - contrary to standard prompting - substantially improves reasoning in compact models. This insight challenges prevailing CoT conventions and reveals new alignment strategies under limited model capacity. Our findings offer a practical blueprint for training deployable, reasoning-rich VLMs suited for mobile and edge applications.
Problem

Research questions and friction points this paper is trying to address.

Video-Language Models
Reasoning Distillation
Video Question Answering
Compact Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Efficient Reasoning Distillation
Synthetic Chain-of-Thought
Difficulty-Aware Fine-Tuning
Compact Video-Language Models
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
M
Mantek Singh
Liverpool John Moores University, Liverpool, England
J
Jeshwanth Challagundla
Carnegie Mellon University, Pittsburgh, USA
S
Siddharth Raina
Meta, Sunnyvale, USA
J
Jasmin Jarsania
University of Texas at Arlington, Texas, USA