How effective are VLMs in assisting humans in inferring the quality of mental models from Multimodal short answers?

πŸ“… 2026-02-10
πŸ›οΈ Conference of the European Chapter of the Association for Computational Linguistics
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of assessing the quality of students’ mental models in multimodal short-answer responses, which requires deep reasoning beyond the capabilities of traditional scoring methods that often fail to capture conceptual understanding. To this end, the authors propose MMGrader, a novel approach that integrates vision-language models (VLMs) with concept graphs to jointly model the semantic content and structural organization of multimodal answers, enabling interpretable and structured evaluation of mental model quality. Evaluated across nine open-source models, MMGrader achieves state-of-the-art performance with an accuracy of 40% and a prediction error of 1.1 points, while its score distributions align closely with human grading trends. This work offers educators an effective AI-driven paradigm for diagnosing collective classroom understanding at scale.
πŸ“ Abstract
STEM Mental models can play a critical role in assessing students'conceptual understanding of a topic. They not only offer insights into what students know but also into how effectively they can apply, relate to, and integrate concepts across various contexts. Thus, students'responses are critical markers of the quality of their understanding and not entities that should be merely graded. However, inferring these mental models from student answers is challenging as it requires deep reasoning skills. We propose MMGrader, an approach that infers the quality of students'mental models from their multimodal responses using concept graphs as an analytical framework. In our evaluation with 9 openly available models, we found that the best-performing models fall short of human-level performance. This is because they only achieved an accuracy of approximately 40%, a prediction error of 1.1 units, and a scoring distribution fairly aligned with human scoring patterns. With improved accuracy, these can be highly effective assistants to teachers in inferring the mental models of their entire classrooms, enabling them to do so efficiently and help improve their pedagogies more effectively by designing targeted help sessions and lectures that strengthen areas where students collectively demonstrate lower proficiency.
Problem

Research questions and friction points this paper is trying to address.

mental models
multimodal answers
conceptual understanding
VLMs
STEM education
Innovation

Methods, ideas, or system contributions that make the work stand out.

mental models
multimodal assessment
concept graphs
VLMs
MMGrader
P
Pritam Sil
Department of Computer Science and Engineering, IIT Bombay, Mumbai, India
D
Durgaprasad Karnam
Center for Educational Technology, IIT Bombay, Mumbai, India
V
Vinay Reddy Venumuddala
School of Management, Mahindra University, Hyderabad, India
P
Pushpak Bhattacharyya
Department of Computer Science and Engineering, IIT Bombay, Mumbai, India