Temporal Self-Distillation: Learning Visual State Tracking in Videos Without Supervision

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出S^3T方法,通过自监督蒸馏学习视频状态跟踪,无需标签或额外教师模型,提高了VSTAT和MVBench上的性能。
📝 Abstract
We introduce S$^3$T (Self-Supervised Self-Distillation over Time), which, to the best of our knowledge, is the first fully self-contained framework for continuous video state tracking. Our method treats temporal sampling density as privileged information, based on the hypothesis that a denser view of the same clip recovers the running state more accurately. This view serves as the teacher, while a sparse-view student with the same weights learns to match its next-token distribution. The model generates its own target, so training requires no labels, separate teacher, or reward signal, and adds no inference cost. On LLaVA-OneVision-2-8B, S$^3$T improves VSTAT accuracy by $+1.74$ as a single model, $+2.38$ with souping, and $+2.70$ with additional vision-encoder adaptation, while prior self-evolving methods leave state tracking largely unchanged. The capability learned from unlabeled synthetic clips transfers to real videos, improving performance by $+7.95$ on VSTAT-YouTube state-tracking questions and $+4.50$ on MVBench Action Count.
Problem

Research questions and friction points this paper is trying to address.

Temporal Self-Distillation
Visual State Tracking
Unsupervised Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Supervised Learning
Temporal Self-Distillation
State Tracking
Unsupervised Video Analysis
Transfer Learning
🔎 Similar Papers
No similar papers found.