Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video Captioning

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决弱监督密集视频字幕中辅助过渡字幕缺乏视觉基础的问题,提出了一种基于VLM的框架SBS,通过生成帧级叙述并检测语义变化来优化过渡事件定位。
📝 Abstract
Weakly-Supervised Dense Video Captioning aims to localize and describe multiple events in untrimmed videos given only an ordered set of event-level captions per video. Recent work synthesizes auxiliary transition captions via LLM to provide additional vision-language alignment, but these captions lack visual grounding and are rigidly assigned to every inter-event gap at a fixed location and duration. To address these, we propose Seeing Before Synthesizing (SBS), a framework that adaptively provides visually grounded linguistic guidance only where warranted. Leveraging a VLM, we generate frame-level narratives for the inter-event gaps and detect transitions from the semantic variation across them. For identified transitions, we then refine inter-event temporal masks by blending the temporal midpoint with the semantic change point and selecting the width that maximizes vision-language alignment. Experiments on ActivityNet Captions and YouCook2 demonstrate state-of-the-art performance in both captioning and localization.
Problem

Research questions and friction points this paper is trying to address.

Weakly-Supervised Dense Video Captioning
Transition Event Discovery
Vision-Language Alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

VLM-Guided
Transition Event Discovery
Visually Grounded Linguistic Guidance
Temporal Masks Refinement
🔎 Similar Papers
Y
Ye-Chan Kim
Hanyang University, South Korea
S
Seunghee Choi
Hanyang University, South Korea
S
SeungJu Cha
Hanyang University, South Korea
S
Si-Woo Kim
Hanyang University, South Korea
H
Hwiseon Kim
Hanyang University, South Korea
H
Hyungee Kim
Hanyang University, South Korea
Dong-Jin Kim
Dong-Jin Kim
Assistant Professor, Hanyang University
Computer VisionMachine LearningNatural Language ProcessingArtificial Intelligence