Cap2Sum: Learning to Summarize Videos by Generating Captions
To address the high cost of manual annotation and limited data scale in video summarization—which severely hinder model generalization—this paper proposes a large-scale training paradigm leveraging dense video captions as weak supervision. Methodologically, it pioneers the use of dense captions instead of human-generated summaries as supervisory signals; incorporates CLIP’s vision-language priors to explicitly recover salient objects missing from captions; and designs a Transformer-based cross-modal generation architecture that integrates weakly supervised learning with zero-shot transfer and cross-dataset fine-tuning strategies. Evaluated on two newly constructed benchmarks—TVSum-Caption and SumMe-Caption—the approach substantially outperforms prior methods, achieving significant improvements in summary quality and cross-domain generalization. This work establishes a viable pathway toward low-cost, large-scale video summarization.