VIP-MINGLE: A Corpus for Videoconference and In-Person Multimodal Interaction in Group Language Engagement

πŸ“… 2026-07-15
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses a critical gap in group multimodal dialogue researchβ€”the lack of datasets encompassing both video conferencing and face-to-face interaction scenarios. To bridge this gap, the authors introduce a cross-scenario multimodal dataset comprising 59 hours of recordings from 32 groups involving 105 participants. The dataset includes synchronized raw audiovisual streams and psychometric measurements, along with processed multimodal features such as voice activity segmentation, facial expression recognition, speech transcripts, and time-aligned annotations. Notably, it is the first to systematically capture paired group interactions across both communication environments, revealing significant differences in multimodal behavior between settings. This resource provides a foundational benchmark for developing robust models of group dialogue that generalize across interaction contexts.
πŸ“ Abstract
Group conversations are a fundamental yet complex form of social interaction central to human cognition and telecommunication technology. While understanding and facilitating these interactions has been a long-standing goal, findings are often isolated within specific in-person or videoconferencing settings due to a scarcity of datasets that bridge the two. We introduce VIP-MINGLE, a multimodal dataset comprising 59 hours of recordings (32 groups, 105 participants), featuring paired within-subject sessions in both settings. The dataset includes raw audio/video, psychometric data, processed multimodal features (e.g., diarized speech, facial expressions, transcriptions), and time-resolved human annotations. Our analysis reveals significant behavioral distribution shifts across multiple modalities between settings, reinforcing the need for a cross-setting corpus. VIP-MINGLE serves as a critical resource for developing robust models of group conversations across settings.
Problem

Research questions and friction points this paper is trying to address.

group conversation
videoconferencing
in-person interaction
multimodal dataset
cross-setting
Innovation

Methods, ideas, or system contributions that make the work stand out.

multimodal dataset
cross-setting interaction
group conversation
videoconference
in-person communication
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Andrew Chang
Andrew Chang
Department of Psychology, New York University
cognitive neurosciencemachine learningauditory perceptioninterpersonal interaction
A
Abhinay K Bodi
New York University, USA
W
Wenxin Deng
New York University, USA
J
Junrui Huang
New York University, USA
V
Venu G Kadamba
New York University, USA
S
Sumanth B H Karanam
New York University, USA
D
Dhiwahar A Kennady
New York University, USA
David Poeppel
David Poeppel
Max Planck Society & NYU
Dustin Freeman
Dustin Freeman
On DIY Sabbatical: dustinfreeman.org/blog/sabbatical-themes/
Human Computer Interaction