ChitraMiti: Benchmarking Visual Grounding and Modality Reliance in Bengali Geometric Reasoning

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过创建ChitraMiti-12.8k和NCTB-500数据集及三阶段评估协议,解决了孟加拉语几何推理中视觉基础与模态依赖性问题。
📝 Abstract
Evaluation of vision-language models (VLMs) for multimodal mathematical reasoning remains limited for low-resource languages and for geometry problems that require reading a diagram and a question together. We introduce ChitraMiti-12.8k, a synthetic benchmark of 12,874 Bengali planar geometry problems paired with structured 15-attribute descriptions, and NCTB-500, a complementary set of 500 diagrams manually extracted from Bengali school textbooks. Using a three-phase protocol that separates diagram-only, diagram-plus-description, and description-only inputs, we show across five open-weight and closed-source VLMs that description-only performance is statistically indistinguishable from diagram-plus-description performance, establishing structured descriptions as a sufficient textual proxy for controlled evaluation. Despite this, models remain poor at cross-modal verification, frequently misled by a swapped spatial relation even when they answer the unmodified item correctly. We further evaluate supervised adaptation on ChitraMiti-12.8k, finding that fine-tuning improves performance on both ChitraMiti-1k and NCTB-500, although a substantial gap to the strongest zero-shot model remains. Together, ChitraMiti-12.8k, NCTB-500, and our evaluation protocol offer a standardized way to study Bengali multimodal geometry reasoning and, more broadly, whether VLMs actually check their text against what they see. Our dataset and code are publicly available on Hugging Face at https://huggingface.co/datasets/RaiyanKhaan/ChitraMiti.
Problem

Research questions and friction points this paper is trying to address.

Bengali
Geometric Reasoning
Vision-Language Models
Multimodal
Cross-modal Verification
Innovation

Methods, ideas, or system contributions that make the work stand out.

multimodal mathematical reasoning
low-resource languages
structured descriptions
cross-modal verification
supervised adaptation
🔎 Similar Papers
K
Khan Raiyan Ibne Reza
North South University
S
Sanjana Aktar Maria
North South University
S
Sumaiya Tabassum Nimi
North South University
Md Adnan Arefeen
Md Adnan Arefeen
Assistant Professor @ NSU, PhD @ University of Missouri, Ex Research Intern @ NEC Labs America
Deep LearningComputer VisionEdge ComputingGenerative AIVideo Analytics