Institution profile

Max Planck Institute for Psycholinguistics

Academic institutioneurope · nl
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Investigating Multimodal Informativity under Different Partner Visibility Conditions in Video-Mediated Dialogue

Aug 09, 2026

This study investigates how interactional visibility modulates the informational contributions of gesture and speech in referential tasks within video-mediated dialogue. By constructing models based on speech transcripts, gesture skeleton sequences, and their multimodal fusion—trained with alignment between representations and referential images—the research systematically evaluates the referential efficacy of each modality under varying visibility conditions. Findings demonstrate that gestures possess independent referential capacity, and multimodal fusion yields the greatest performance gains when speech is highly ambiguous. Moreover, human behavioral analyses reveal that visibility not only regulates the informativeness of gestural production but also elicits cross-turn verbal coordination, underscoring the critical role of pragmatic factors in multimodal reference.

0 citationsRead paper

You Prefer This One, I Prefer Yours: Using Reference Words is Harder Than Vocabulary Words for Humans and Multimodal Language Models

May 29, 2025

This study investigates cognitive limitations of multimodal language models (MLMs) in referential comprehension—specifically, possessive pronouns (e.g., “my”/“your”) and demonstratives (e.g., “this”/“that”)—which are frequent yet challenging in everyday communication. Method: We construct the first cognitively grounded hierarchy of referential difficulty, and systematically evaluate seven state-of-the-art MLMs via human behavioral benchmarks, cross-modal coreference resolution tasks, and prompt-engineering interventions. Contribution/Results: MLMs significantly underperform humans on demonstrative comprehension; prompt engineering yields only marginal gains for possessives and fails to close the demonstrative gap. We identify a fundamental structural deficit: insufficient perspective-taking and pragmatic inference capabilities. This work provides the first empirical characterization of MLMs’ critical shortcomings in social cognition, establishing a new evaluation benchmark and theoretical foundation for advancing multimodal language understanding.

0 citationsRead paper
Recent publications

Latest Papers

Investigating Multimodal Informativity under Different Partner Visibility Conditions in Video-Mediated Dialogue

Aug 09, 2026

This study investigates how interactional visibility modulates the informational contributions of gesture and speech in referential tasks within video-mediated dialogue. By constructing models based on speech transcripts, gesture skeleton sequences, and their multimodal fusion—trained with alignment between representations and referential images—the research systematically evaluates the referential efficacy of each modality under varying visibility conditions. Findings demonstrate that gestures possess independent referential capacity, and multimodal fusion yields the greatest performance gains when speech is highly ambiguous. Moreover, human behavioral analyses reveal that visibility not only regulates the informativeness of gestural production but also elicits cross-turn verbal coordination, underscoring the critical role of pragmatic factors in multimodal reference.

0 citationsRead paper

You Prefer This One, I Prefer Yours: Using Reference Words is Harder Than Vocabulary Words for Humans and Multimodal Language Models

May 29, 2025

This study investigates cognitive limitations of multimodal language models (MLMs) in referential comprehension—specifically, possessive pronouns (e.g., “my”/“your”) and demonstratives (e.g., “this”/“that”)—which are frequent yet challenging in everyday communication. Method: We construct the first cognitively grounded hierarchy of referential difficulty, and systematically evaluate seven state-of-the-art MLMs via human behavioral benchmarks, cross-modal coreference resolution tasks, and prompt-engineering interventions. Contribution/Results: MLMs significantly underperform humans on demonstrative comprehension; prompt engineering yields only marginal gains for possessives and fails to close the demonstrative gap. We identify a fundamental structural deficit: insufficient perspective-taking and pragmatic inference capabilities. This work provides the first empirical characterization of MLMs’ critical shortcomings in social cognition, establishing a new evaluation benchmark and theoretical foundation for advancing multimodal language understanding.

0 citationsRead paper