MathScape: Evaluating MLLMs in multimodal Math Scenarios through a Hierarchical Benchmark
Existing mathematical reasoning benchmarks heavily rely on synthetic images, failing to capture the complexity of real-world multimodal reasoning involving photographs and textual mathematics. Method: We introduce MathScape—the first hierarchical multimodal benchmark for photorealistic mathematical problems—featuring a novel “scene–semantics–task” three-level taxonomy that systematically integrates authentic images with formal mathematical semantics, thereby addressing the longstanding gap in joint vision-language mathematical reasoning evaluation. Contribution/Results: Leveraging 11 state-of-the-art multimodal large language models (MLLMs), we conduct dual-track evaluation assessing both theoretical understanding and practical application. Empirical results reveal that even top-performing models achieve sub-50% average accuracy, exposing critical weaknesses in cross-modal alignment, symbolic parsing, and multi-step reasoning. MathScape establishes a new, high-challenge, fine-grained, and interpretable evaluation paradigm for multimodal mathematical reasoning.