MathAgent: Leveraging a Mixture-of-Math-Agent Framework for Real-World Multimodal Mathematical Error Detection
Accurately identifying errors in K–12 multimodal math assignments—comprising both handwritten or typeset text and diagrams—remains challenging, as current multimodal large language models (MLLMs) lack the capability to jointly reason over image and text modalities and precisely localize and attribute solution-step errors. Method: This paper proposes the first three-stage mathematical agent hybrid framework: (1) image-text consistency verification, (2) visual-semantic parsing, and (3) cross-modal error integration analysis—explicitly modeling multimodal associations between problem statements and solution steps. The framework integrates vision understanding, symbolic logical reasoning, and pedagogically grounded constraints within a specialized collaborative agent architecture. Contribution/Results: Evaluated on real-world educational data, the framework improves step-level error detection accuracy by 5% and error-type classification accuracy by 3%. It has been deployed at scale across a platform serving over one million students, achieving 90% user satisfaction and substantially reducing manual review overhead.