Read the Room, Read the Image: Understanding Indirect Speech Acts in Multimodal Visual Contexts

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过构建READI基准,结合视觉和对话信息解决间接言语行为理解问题,尤其针对高语境语言如韩语。
📝 Abstract
Indirect speech acts (ISAs) require pragmatic reasoning over context, as directive intent can- not be inferred from surface form alone. Prior text-based studies and existing multimodal benchmarks largely overlook this requirement, focusing instead on explicitly encoded context or perceptual recognition, and thus underex- plore context-dependent pragmatic understand- ing, particularly in high-context languages such as Korean. We introduce READI, a multimodal benchmark for evaluating ISA understanding through integrated reasoning over visual con- text and dialogue. READI models graded in- directness grounded in pragmatic theory and formulates the task as vision-based pragmatic question answering (V-PQA), supporting cross- lingual evaluation in English and Korean. Ex- periments show that even state-of-the-art multi- modal models struggle with visually grounded indirect speech acts, with performance declin- ing as indirectness increases, underscoring the need for benchmarks that explicitly target con- textual pragmatic reasoning.
Problem

Research questions and friction points this paper is trying to address.

Indirect Speech Acts
Multimodal Context
Pragmatic Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

multimodal benchmark
indirect speech acts
visual context
pragmatic reasoning
cross-lingual evaluation
💼 Related Jobs
No related jobs found.
Jaehee Kim
Jaehee Kim
Assistant Professor at Cornell University, Department of Computational Biology
J
Ji Hoon Chung
Department of Korean Language and Literature, Yonsei University
S
Seoyoon Park
Interdisciplinary Graduate Program of Linguistics and Informatics, Yonsei University
U
Unsol Kim
LG AI Research
K
Kyungwon Park
Department of Artificial Intelligence, Yonsei University
J
Ji Hak Kim
Interdisciplinary Graduate Program of Linguistics and Informatics, Yonsei University
Y
Yi-Jun Chen
Department of Korean Language and Literature, Yonsei University
H
Hansaem Kim
Interdisciplinary Graduate Program of Linguistics and Informatics, Yonsei University