Multi-Branch Policy Optimization for Multimodal Large Language Models

๐Ÿ“… 2026-08-05
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the challenge that multimodal large language models in reinforcement learning suffer from degraded learning signals due to trajectory-level unified advantage estimation under high visual perceptual uncertainty. To mitigate this issue, the authors propose Multibranch Policy Optimization (MBPO), a novel framework that introduces a tree-structured reasoning mechanism into multimodal reinforcement learning for the first time. Specifically, MBPO constructs a reasoning tree at the visionโ€“language decision boundary, where sibling branches explore alternative visual hypotheses and enable segment-level credit assignment through relative advantages across branches. Additionally, a temporal replay buffer is incorporated to enhance sample efficiency and training stability. Experimental results demonstrate that MBPO significantly outperforms existing baselines on multiple multimodal reasoning benchmarks and effectively alleviates credit assignment failure under conditions of high perceptual uncertainty.
๐Ÿ“ Abstract
Group-based reinforcement learning methods for multimodal large language models typically rely on trajectory-level credit assignment that applies a single advantage to all tokens in a response. However, multimodal reasoning involves substantially higher perceptual uncertainty than text-only settings, where the model must repeatedly re-examine visual information to verify intermediate interpretations, and different visual groundings can lead to divergent reasoning paths, making such uniform credit assignment particularly inadequate and causing relative advantages to progressively degenerate toward zero. To address these challenges, we propose Multi-Branch Policy Optimization (MBPO), a tree-based framework that constructs reasoning trees at vision-language decision boundaries, enabling sibling branches to explore diverse visual hypotheses and assigning segment-level credit through branch-relative advantages. We further introduce a temporal replay buffer to reuse informative segments while controlling policy staleness. Experiments on several multimodal reasoning benchmarks show that MBPO outperforms representative baselines, improving both learning signal quality and optimization efficiency. The code is publicly available at https://github.com/ShuaiLyu0110/MBPO.
Problem

Research questions and friction points this paper is trying to address.

multimodal reasoning
credit assignment
perceptual uncertainty
visual grounding
reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Branch Policy Optimization
reasoning trees
segment-level credit assignment
branch-relative advantages
temporal replay buffer
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
S
Shuai Lyu
Beijing University of Posts and Telecommunications
Y
Yuning Gong
Sichuan University
R
Ruiling Gao
Shanghai University
X
Xiaoran Shang
Beijing University of Posts and Telecommunications
Zhonghong Ou
Zhonghong Ou
School of Computer Science, Beijing University of Posts and Telecommunications (BUPT), China
Computer VisionDeep LearningMachine LearningBig Data Analytics
P
Ping Zong
Beijing University of Posts and Telecommunications
Yifan Zhu
Yifan Zhu
Beijing University of Posts and Telecommunications
PEFT of LLMsGraph RAGGraph mining
Y
Yuan Sun
Sichuan University
Yang Qin
Yang Qin
Huawei Technologies Ltd.
P
Peng Hu
Sichuan University