CausalChapter: Improving Long-Video Chaptering with Interventional Dependency Modeling

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出CausalChapter框架,通过轻量级干预解决长视频自动分章中的边界误差传播和跨章节上下文碎片化问题,提高边界定位、章节描述质量和连贯性。
📝 Abstract
Long-form instructional videos require automatic chaptering to support browsing, navigation, and knowledge access. Recent long-context language models can perform chaptering from textualized video inputs, but they remain costly and brittle for content-dense lecture videos with long transcripts, smooth topic transitions, and detailed chapter outputs. A scalable segment-then-caption paradigm reduces this cost, but introduces two new challenges: boundary error propagation and fragmented cross-chapter context. We propose \textbf{CausalChapter}, an intervention-inspired framework for long-video chaptering that estimates prediction-level influence through lightweight masking and removal interventions. For boundary localization, our Local Dependency Shift module detects drops in predictive dependency between adjacent temporal windows; for chapter description generation, our Cross-Segment Support Selection module reranks historical contexts according to their support for the current prediction. Experiments on long-video chaptering benchmarks show that CausalChapter improves boundary localization, chapter description quality, and cross-chapter coherence.
Problem

Research questions and friction points this paper is trying to address.

long-video chaptering
instructional videos
boundary error propagation
cross-chapter context
Innovation

Methods, ideas, or system contributions that make the work stand out.

Interventional Dependency Modeling
Local Dependency Shift
Cross-Segment Support Selection
X
Xinran Duan
School of Artificial Intelligence, Beijing Normal University; Beijing Key Laboratory of Artificial Intelligence for Education; Engineering Research Center of Intelligent Technology and Educational Application, Ministry of Education
G
Guozhang Li
School of Artificial Intelligence, Beijing Normal University; Beijing Key Laboratory of Artificial Intelligence for Education; Engineering Research Center of Intelligent Technology and Educational Application, Ministry of Education
Yaoyao Zhong
Yaoyao Zhong
Beijing Normal University
Computer VisionMultimediaAdversarial Robustness
Mei Wang
Mei Wang
Beijing Normal University
face recognitionfairness in AIdomain adaptation
Lizhi Wang
Lizhi Wang
Beijing Normal University
Computational PhotographyComputer VisionMultimedia ComputingImage Processing
Hua Huang
Hua Huang
Beijing Normal University
Visual ComputingComputer GraphicsComputational Photography