MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决教育场景中大型视觉-语言模型评估不足的问题,通过引入MUSE基准测试,涵盖艺术图像理解的多个任务,以提高模型在情感解读和组合推理等方面的能力。
📝 Abstract
Large vision-language models have achieved remarkable progress in multi-modal understanding, yet their capabilities in educational settings remain insufficiently evaluated. In AI-assisted language learning, models must interpret artistic imagery, understand its semantic, affective, and cultural content, and reason about visual context to support meaningful interaction. However, existing benchmarks primarily focus on real-world images or domain-specific educational reasoning, providing limited coverage of artistic educational content. To address this gap, we introduce MUSE, a benchmark for evaluating large vision-language models on artistic image understanding in situated educational applications. MUSE decouples image annotation from question generation, enabling diverse tasks with controllable difficulty while reducing annotation effort. It comprises twelve tasks spanning visual perception, semantic and affective interpretation, culture understanding, and compositional reasoning, together with diverse artistic images deliberately curated to center Singaporean and Southeast Asian multicultural contexts alongside Western art traditions, covering multiple themes and difficulty levels. Evaluation of open-source and proprietary models reveals substantial disparities across capability dimensions, particularly in affective interpretation and compositional reasoning. Our analysis further identifies common failure modes and key challenges for developing trustworthy multi-modal models for education. We hope MUSE will serve as a standardized benchmark for advancing multi-modal understanding in situated educational applications.
Problem

Research questions and friction points this paper is trying to address.

large vision-language models
multi-modal understanding
educational settings
artistic imagery
semantic and affective interpretation
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-modal understanding
vision-language models
educational applications
artistic image understanding
benchmark
🔎 Similar Papers
L
Luyao Zhu
AI Singapore, National University of Singapore, Singapore
X
Xun Wei Yee
AI Singapore, National University of Singapore, Singapore
W
Wei Li
Institute of Advanced Intelligence and Computing, A*STAR
M
Mun Thye Mak
AI Singapore, National University of Singapore, Singapore
Wee Siong Ng
Wee Siong Ng
Institute for Infocomm Research
Machine Learning & AIBig DataDistributed Systems