LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决多模态预训练数据需求,通过收集和处理1.3亿个视频URL生成LAION-BVD数据集,采用内容感知场景检测和合成字幕技术,支持视频、音频及图像模态学习。
📝 Abstract
We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a total duration of 10 million hours. The dataset is designed for multimodal pre-training across the video, audio, and image modalities. Using content-aware scene detection, we extract clips for which we synthetically generate video and audio captions. Models trained on these data achieve competitive performance on standard video-text and audio-text benchmarks, with consistent improvements as training or model scale increases. Additionally, we explore video frames as an alternative source of image-text data by extracting scene-changing frames. These frames exhibit a visual distribution distinct from standard web image corpora, and models trained on this dataset achieve strong image-text retrieval performance. We release LAION-BVD to the research community. It significantly expands open access to multimodal videos at an unprecedented scale.
Problem

Research questions and friction points this paper is trying to address.

multimodal pre-training
large-scale dataset
video dataset
audio-text
image-text
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Pre-training
Content-aware Scene Detection
Synthetic Caption Generation
🔎 Similar Papers