VidaForge: Open Research Infrastructure for Video Pretraining Data Recipes

📅 2026-09-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决视频预训练数据管道封闭问题,提出VIDAFORGE,一个可执行五阶段工作流的开放研究基础设施,通过对比不同数据配方对模型性能的影响来优化视频预训练。
📝 Abstract
Video foundation models increasingly rely on large-scale pretraining data, yet the end-to-end data pipelines behind them remain largely closed and difficult to inspect or reuse. Researchers seeking to understand how video data recipes affect model pretraining often need to build substantial infrastructure before testing even a focused hypothesis. We present VIDAFORGE, an open research infrastructure that represents a video data recipe as an executable five-stage workflow from raw videos to training datasets. A decision in this workflow can be varied to construct alternative datasets while preserving how every sample was produced. To demon strate this research workflow, we compare data recipes with different coverage and quality in early from-scratch pretraining of Wan 2.1 and V-JEPA 2.1. Across both learning objectives, the broader-coverage recipe achieves the highest downstream benchmark scores, while loss-based evaluation favors different recipes. This study demonstrates how VidaForge connects data-recipe choices to downstream model performance. We further release VIDAFORGE-3M, containing 3.14 million scene level clips totaling 6,475 hours, with fine-grained annotations and curation signals for video data-recipe research.
Problem

Research questions and friction points this paper is trying to address.

video foundation models
pretraining data
data pipelines
research infrastructure
Innovation

Methods, ideas, or system contributions that make the work stand out.

VIDAFORGE
video data recipe
pretraining
workflow
downstream performance
🔎 Similar Papers
2024-02-20International Conference on Machine LearningCitations: 30
Yan Ma
Yan Ma
Shanghai Jiao Tong University, Generative Artificial Intelligence Research Lab (GAIR)
J
Jiadi Su
Fudan University, Generative Artificial Intelligence Research Lab (GAIR)
Z
Zhulin Hu
Shanghai Jiao Tong University, Shanghai Innovation Institute, Generative Artificial Intelligence Research Lab (GAIR)
Ethan Chern
Ethan Chern
Shanghai Jiao Tong University
Machine LearningNatural Language ProcessingArtificial Intelligence
L
Linhao Zhang
Shanghai University, Generative Artificial Intelligence Research Lab (GAIR)
T
Tiantian Mi
Shanghai Innovation Institute, Fudan University, Generative Artificial Intelligence Research Lab (GAIR)
Pengfei Liu
Pengfei Liu
Associate professor at Shanghai Jiao Tong University
LLM