🤖 AI Summary
This study investigates how interactional visibility modulates the informational contributions of gesture and speech in referential tasks within video-mediated dialogue. By constructing models based on speech transcripts, gesture skeleton sequences, and their multimodal fusion—trained with alignment between representations and referential images—the research systematically evaluates the referential efficacy of each modality under varying visibility conditions. Findings demonstrate that gestures possess independent referential capacity, and multimodal fusion yields the greatest performance gains when speech is highly ambiguous. Moreover, human behavioral analyses reveal that visibility not only regulates the informativeness of gestural production but also elicits cross-turn verbal coordination, underscoring the critical role of pragmatic factors in multimodal reference.
📝 Abstract
Situated language use is multimodal and embodied. For example, gestures can carry information that is absent or underspecified in the speech signal, yet dialogue models typically rely on transcripts alone. We study how much referential information gestures and their combination with speech carry in multimodal dialogue under different partner visibility conditions. % We build models that identify the intended referent in a video-mediated referential communication game based on either the speech transcript, the skeletal representation of gesture, or both modalities. Our results show that gesture alone is predictive of the intended referent and that multimodal fusion is most beneficial when the transcript-based model is uncertain. Training-only alignment of learned representations with the referent image further improves the fusion model performance. % In a comparison with human interaction data, we further see pragmatic effects of interlocutor visibility on gesture production and informativeness as well as an entrainment effect in speech and multimodal, but not gesture, performance across rounds of repeated interaction. We thus make contributions to the technical modelling of multimodal information in human dialogue and the analysis of human interaction data via trained model representations.