UniCAR-RL: Seeing Better before Thinking Deeper in Visual Mathematics

📅 2026-09-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对多模态大语言模型在数学视觉推理中的感知与逻辑错误,提出UniCAR-RL框架,通过解耦和优化感知与推理能力来提高模型性能。
📝 Abstract
Multimodal Large Language Models (MLLMs) often struggle with complex mathematical visual reasoning primarily due to a lack of fine-grained perception, causing initial visual hallucinations to directly trigger cascading reasoning failures. In traditional end-to-end reinforcement learning (RL), sparse rewards fail to decouple perceptual hallucinations from logical missteps, hindering targeted perception optimization. Alternatively, fine-tuning with perception-enhanced CoT data incurs high costs and hallucinations. In this paper, we address these challenges by proposing UniCAR-RL, an annotation-free RL framework. By explicitly decoupling the optimization of perception and reasoning during the training process, it achieves isolation and optimization of both capabilities. Specifically, UniCAR-RL consists of three synergistic branches: 1) a Caption-RL branch that optimizes perception capabilities through verifier-guided reasoning validation; 2) a Reasoning-RL branch that performs logical reasoning based on a gold image description to halt cascading errors; 3) a QA-RL branch that retains native end-to-end alignment to ensure robust question-answering performance. Experiments show that UniCAR-RL substantially improves MLLMs' mathematical and visual reasoning using only raw short-answer data. Furthermore, it demonstrates strong generalization across diverse architectures and scales.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
visual reasoning
perceptual hallucinations
reinforcement learning
fine-grained perception
Innovation

Methods, ideas, or system contributions that make the work stand out.

UniCAR-RL
perception and reasoning decoupling
reinforcement learning
visual mathematics
multimodal large language models
🔎 Similar Papers
No similar papers found.