Rate-Coding Bundle Memory: A Unified Model of Memory and Control for Symbolic Computation in the Brain
该研究提出了一种基于率编码和捆绑记忆系统的混合模型RCBM,旨在结合连接主义和符号系统的优势以解释多种认知现象。
该研究提出了一种基于率编码和捆绑记忆系统的混合模型RCBM,旨在结合连接主义和符号系统的优势以解释多种认知现象。
This study investigates how interactional visibility modulates the informational contributions of gesture and speech in referential tasks within video-mediated dialogue. By constructing models based on speech transcripts, gesture skeleton sequences, and their multimodal fusion—trained with alignment between representations and referential images—the research systematically evaluates the referential efficacy of each modality under varying visibility conditions. Findings demonstrate that gestures possess independent referential capacity, and multimodal fusion yields the greatest performance gains when speech is highly ambiguous. Moreover, human behavioral analyses reveal that visibility not only regulates the informativeness of gestural production but also elicits cross-turn verbal coordination, underscoring the critical role of pragmatic factors in multimodal reference.
This study investigates cognitive limitations of multimodal language models (MLMs) in referential comprehension—specifically, possessive pronouns (e.g., “my”/“your”) and demonstratives (e.g., “this”/“that”)—which are frequent yet challenging in everyday communication. Method: We construct the first cognitively grounded hierarchy of referential difficulty, and systematically evaluate seven state-of-the-art MLMs via human behavioral benchmarks, cross-modal coreference resolution tasks, and prompt-engineering interventions. Contribution/Results: MLMs significantly underperform humans on demonstrative comprehension; prompt engineering yields only marginal gains for possessives and fails to close the demonstrative gap. We identify a fundamental structural deficit: insufficient perspective-taking and pragmatic inference capabilities. This work provides the first empirical characterization of MLMs’ critical shortcomings in social cognition, establishing a new evaluation benchmark and theoretical foundation for advancing multimodal language understanding.
该研究提出了一种基于率编码和捆绑记忆系统的混合模型RCBM,旨在结合连接主义和符号系统的优势以解释多种认知现象。
This study investigates how interactional visibility modulates the informational contributions of gesture and speech in referential tasks within video-mediated dialogue. By constructing models based on speech transcripts, gesture skeleton sequences, and their multimodal fusion—trained with alignment between representations and referential images—the research systematically evaluates the referential efficacy of each modality under varying visibility conditions. Findings demonstrate that gestures possess independent referential capacity, and multimodal fusion yields the greatest performance gains when speech is highly ambiguous. Moreover, human behavioral analyses reveal that visibility not only regulates the informativeness of gestural production but also elicits cross-turn verbal coordination, underscoring the critical role of pragmatic factors in multimodal reference.
This study investigates cognitive limitations of multimodal language models (MLMs) in referential comprehension—specifically, possessive pronouns (e.g., “my”/“your”) and demonstratives (e.g., “this”/“that”)—which are frequent yet challenging in everyday communication. Method: We construct the first cognitively grounded hierarchy of referential difficulty, and systematically evaluate seven state-of-the-art MLMs via human behavioral benchmarks, cross-modal coreference resolution tasks, and prompt-engineering interventions. Contribution/Results: MLMs significantly underperform humans on demonstrative comprehension; prompt engineering yields only marginal gains for possessives and fails to close the demonstrative gap. We identify a fundamental structural deficit: insufficient perspective-taking and pragmatic inference capabilities. This work provides the first empirical characterization of MLMs’ critical shortcomings in social cognition, establishing a new evaluation benchmark and theoretical foundation for advancing multimodal language understanding.