Improving Generalization Robustness of Multimodal RLVR

📅 2026-08-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the fragility of multimodal reinforcement learning with verifiable rewards (RLVR) under prompt format variations, which undermines reliability in high-stakes settings such as medical visual question answering (VQA). To enhance robustness, the authors propose a dynamic triplet reward mechanism that disentangles prompt format from semantic content, coupled with a strategy consistency regularizer based on adversarial perturbations in the embedding space. This regularizer enforces invariant model responses across semantically equivalent but syntactically diverse prompts. The proposed approach substantially improves robust generalization, achieving an average accuracy drop of no more than 1% under stress testing—significantly outperforming GRPO, which exhibits a ~3% decline—and demonstrates state-of-the-art performance in dynamic evaluation scenarios.
📝 Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) makes Multimodal Large Language Models more accurate, but the gains are brittle: simply paraphrasing a question or changing the prompt template can degrade them, which challenges reliable deployment in high-stakes scenarios like medical VQA. We trace this to two issues of the standard RL objective. First, the binary verifier conflates format with content, so the reward signal cannot tell a wrong answer apart from a misformatted one. Second, the training distribution covers only a thin slice of the real-world prompts that the model might meet at deployment, so policies that perform well on the training distribution can behave differently under unseen prompts during test. Both failures call for a robust post-training method that helps the policy cover a broader distribution of semantically equivalent prompts, and we identify two measures that help achieve this objective: separating format from semantics in the reward, and applying policy invariance across perturbed prompts with equivalent semantics. We therefore propose Prompt-Invariant RLVR (PIRL), consisting of a dynamic trinary reward and a consistency regularizer based on an embedding-space adversary. Under stress testing, PIRL's average accuracy on benchmarks drops by only $\le 1\%$, where GRPO drops ~3%. On dynamic evaluation, PIRL also achieves the smallest performance drop.
Problem

Research questions and friction points this paper is trying to address.

generalization robustness
multimodal RLVR
prompt sensitivity
reward design
distribution shift
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prompt-Invariant RLVR
trinary reward
consistency regularization
semantic invariance
robust generalization
🔎 Similar Papers
No similar papers found.
P
Pengfei Zhou
National University of Singapore, HPC-AI Lab
Zhiwei Tang
Zhiwei Tang
Professor, University of Electronic Science and Technology of China
X
Xiaopeng Peng
Rochester Institute of Technology
C
Chenrui Zhou
National University of Singapore, HPC-AI Lab
L
Lama Moukheiber
Georgia Institute of Technology
Y
Yixing Ma
University of California, Berkeley
B
Bin Xu
InfRec, Cardinal AI Lab
Jiajun Song
Jiajun Song
Michigan technological University
Wave Energy Converter
Z
Zhenglin Wan
National University of Singapore, HPC-AI Lab
Wangbo Zhao
Wangbo Zhao
National University of Singapore
Efficient Deep LearningDynamic Neural NetworkMultimodal Model
J
Jiasheng Tang
DAMO Academy, Alibaba Group; Hupan Lab
Bohan Zhuang
Bohan Zhuang
Zhejiang University
Efficient AIMLSys
Fan Wang
Fan Wang
Alibaba DAMO Academy
Computer VisionMachine Learning
Yang You
Yang You
Postdoc, Stanford University
3D visioncomputer graphicscomputational geometry