Func-R1: Incentivizing Mathematical Function Reasoning in Multimodal Large Language Models

📅 2026-09-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了多模态大语言模型在数学函数推理中忽视视觉线索的问题,通过Func-R1方法结合精确视觉感知与逻辑推理来改善。
📝 Abstract
Performing deliberate mathematical reasoning in visual contexts is a hallmark of advanced Multimodal Large Language Models (MLLMs) and requires a sophisticated synthesis of perceptual grounding and symbolic logic. However, in the realm of mathematical functions, our investigation reveals a critical modality interference phenomenon: even advanced models, while performing textual computational reasoning, tend to disregard or misinterpret essential visual cues. To address this challenge, we propose Func-R1, which synergistically harmonizes precise visual perception and rigorous logical reasoning. Concretely, built upon an explicitly decoupled architecture, we employ a hierarchical post-training framework to progressively identify critical visual evidence and conduct in-depth theoretical reasoning. Furthermore, the Perception-Aligned Theoretic Optimization (PATO) strategy is proposed to steer policy updating towards internalizing fundamental theoretical properties while dynamically rectifying heterogeneous visual information throughout the reasoning process. Extensive experiments across diverse benchmarks demonstrate that Func-R1 delivers the optimal performance among open-source MLLMs, even surpassing GPT-5 with an 8.4% improvement on MathVerse's function-oriented tasks.
Problem

Research questions and friction points this paper is trying to address.

mathematical function reasoning
multimodal large language models
visual cues
modality interference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Func-R1
hierarchical post-training framework
Perception-Aligned Theoretic Optimization (PATO)
multimodal large language models
mathematical function reasoning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.