Which one is banana man? Evaluating vision-language models in multi-turn pragmatic interpretation

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过迭代引用游戏测试了视觉-语言模型在多轮对话中的语用理解能力,发现这些模型虽能利用先前上下文但难以有效构建相关上下文。
📝 Abstract
Flexible adaptation to context and shared pragmatic intuitions contribute to smooth human conversation. Iterated reference games---in which players repeatedly pick out novel referents using language---present a test case for agents' ability to perform context-sensitive pragmatic reasoning in multi-turn linguistic environments. We tested humans and vision--language models on their ability to identify the intended meaning of descriptions produced in iterated reference games, varying the provided context in terms of amount, order, and relevance. While humans performed well consistently, the models we evaluated could make use of prior context to interpret humans' referring expressions, but they struggled to build up the relevant context to interpret those expressions effectively. Our results suggest that the models we evaluated lack core skills needed for efficient linguistic collaboration.
Problem

Research questions and friction points this paper is trying to address.

vision-language models
pragmatic interpretation
iterated reference games
context-sensitive reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

iterated reference games
context-sensitive pragmatic reasoning
vision-language models
🔎 Similar Papers