Training-Free Temporal Abstraction for General Video Understanding

๐Ÿ“… 2026-08-28
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡ๆๅ‡บSTITCH๏ผŒไธ€็งๆ— ้œ€่ฎญ็ปƒ็š„ๆ–นๆณ•๏ผŒ้€š่ฟ‡้ข„่ฎญ็ปƒ็š„่ง†้ข‘-ๆ–‡ๆœฌๆจกๅž‹ๅฐ†่ง†้ข‘ๅˆ†ๅ‰ฒๆˆๆœ‰ๆ„ไน‰็š„ๆ—ถ้—ดๅ—๏ผŒไปฅๆ”ฏๆŒๅคš็ง่ง†้ข‘็†่งฃไปปๅŠกใ€‚
๐Ÿ“ Abstract
Videos are expensive to analyze frame by frame, yet many video understanding tasks depend on knowing where relevant moments occur. A system may need to find when an action changes, locate the segment described by a sentence, or choose a few frames for a vision-language model. Existing methods often solve these problems separately, using task-specific training data or specialized architectures. We study whether a pretrained video-text model can provide enough temporal structure to support several of these tasks at once. We present STITCH, a training-free method that divides a video into semantically meaningful temporal chunks. STITCH embeds short video windows with a frozen video-text backbone and detects changes in the resulting embedding sequence. These chunks are computed once per video and reused across tasks. We evaluate STITCH on generic event boundary detection, language-based moment retrieval, and frame selection for long-video VLM reasoning. Across all three settings, STITCH remains competitive with more specialized methods while requiring no task-specific training, with especially clear gains when only a small number of frames or tokens can be processed. These results suggest that reusable temporal abstraction is a promising direction for general video understanding, allowing dense video streams to be converted once into semantic units that can be localized, retrieved, sampled, or reasoned over by downstream systems.
Problem

Research questions and friction points this paper is trying to address.

video understanding
temporal abstraction
pretrained model
task-specific training
semantic chunks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-Free
Temporal Abstraction
Pretrained Video-Text Model
Semantic Chunks
General Video Understanding
๐Ÿ”Ž Similar Papers
2024-02-20International Conference on Machine LearningCitations: 30
๐Ÿ’ผ Related Jobs
No related jobs found.
E
Etienne Casanova
California Institute of Technology
S
Sevan Brodjian
California Institute of Technology
Pietro Perona
Pietro Perona
California Institute of Technology
Computer VisionMachine LearningApplied MathematicsNeurosciencePsychology