InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos

📅 2026-08-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing audio-visual generation models struggle to produce natural and causally coherent multimodal responses due to the absence of real-world user interaction supervision data. This work proposes and constructs InteracVid, the first large-scale open-source interactive audio-visual dataset, by extracting context–stimulus–response triplets from live-stream videos across diverse scenarios including dialogue, physical manipulation, and demonstrations. Leveraging a metadata-aware mining pipeline, the authors automatically curate 454K high-quality triplets from over 59K videos amidst noisy live-stream content, explicitly encoding causal structure, temporal completeness, and naturalness. Experiments demonstrate that fine-tuning on this dataset substantially improves interactive planning and response generation quality, with consistent gains observed across both human evaluation and automatic metrics on a benchmark of 100 real-world chat queries.
📝 Abstract
Large language models have made text the default medium for human--AI interaction, buttext alone cannot express the full range of responses required by multimodal assistants,avatars, and embodied agents. While recent audio-video generative models can synthesizehigh-fidelity synchronized content, existing supervision is largely \emph{descriptive}:models are trained to render captions rather than to produce audio-visual responsescaused by external user interactions. We introduce \textbf{InteracVid}, \emph{the firstopen-source large-scale dataset that addresses this missing supervision}, so that everysample couples a preceding audio-visual context and an external stimulus with the realinteractive response that follows. We design a metadata-aware pipeline that extractsinteractive clips from long, noisy livestreams, yielding over \textbf{454K}context-query-response triplets from more than \textbf{59K} livestream videos andspanning conversation-centered, object-centric, procedural, embodied, and screen-basedscenarios. A ten-rater human study confirms that the extracted interactions are causal,natural, and temporally complete for both genuine and reconstructed queries. On aheld-out benchmark of \textbf{100} genuine live-chat queries, fine-tuning on InteracVidimproves both interaction planning and audio-video response generation, and anindependent human evaluation reproduces the system ranking and the conclusions obtainedwith our automatic judge. These results highlight interaction-structured data as acritical foundation for interactive multimodal generation.
Problem

Research questions and friction points this paper is trying to address.

interactive multimodal generation
audio-visual response
causal supervision
live-chat videos
multimodal interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

interactive multimodal generation
audio-visual response dataset
causal interaction supervision
live-chat video mining
context-query-response triplets
🔎 Similar Papers
2024-06-09Annual Meeting of the Association for Computational LinguisticsCitations: 13
C
Chi Zhang
College of AI, Tsinghua University
H
Haoyang Shi
College of AI, Tsinghua University
Y
Yueyi Liu
College of AI, Tsinghua University
Z
Zhaokun Yan
College of AI, Tsinghua University
Y
Yishu Yin
College of AI, Tsinghua University
Yuhang Wu
Yuhang Wu
Universitat Pompeu Fabra
Miao Liu
Miao Liu
Assistant Professor at Tsinghua University, College of AI
Computer VisionDeep LearningAugmented RealityGenerative Model