Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs

📅 2026-08-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the inconsistent response behaviors exhibited by multimodal large language models with hybrid reasoning modes—specifically, discrepancies between “thinking” and “non-thinking” inference pathways—where correctness alone is insufficient to assess response quality. The study introduces the first systematic formulation of the response pattern alignment problem and proposes PatternEval, a benchmark encompassing four representative failure modes. To align user-visible responses across reasoning modes, the authors develop PatternRM, a reward model, and PatternRL, a reinforcement learning framework that incorporates pattern-specific penalties. Experiments on Qwen3-VL-4B and Qwen3-VL-8B demonstrate that PatternRL substantially reduces failure rates in non-thinking mode, effectively mitigates cross-mode mismatch, and achieves these improvements with negligible impact on overall task performance.
📝 Abstract
Hybrid-thinking multimodal large language models (MLLMs) allow a single model to alternate between deliberative thinking and latency-efficient non-thinking inference. Although these modes differ in reasoning budget, their delivered responses should satisfy the same user-facing standard. Correctness alone may not characterize this response quality; we therefore evaluate task accuracy and response-pattern failures as complementary outcomes. We study this gap through \textbf{response-pattern alignment}: whether thinking and non-thinking interfaces preserve acceptable final-response behavior. We introduce \textbf{PatternEval}, a failure-enriched diagnostic benchmark comprising 2,415 multimodal prompts spanning visual perception and grounding, structured image understanding, and multimodal knowledge reasoning. PatternEval tests four recurrent failures: chain-of-thought leakage, response repetition, logical contradiction, and performative reasoning. Response-pattern failures are widespread across models from different providers, with non-thinking inference exhibiting substantially higher failure rates and thereby creating systematic misalignment between thinking and non-thinking interfaces. Motivated by this diagnosis, we develop \textbf{PatternRM}, a response-level reward model, and \textbf{PatternRL}, which introduces pattern-specific penalties during reinforcement learning. Experiments on Qwen3-VL-4B and Qwen3-VL-8B show that incorporating pattern-specific penalties into reinforcement learning can mitigate cross-mode misalignment while incurring a marginal task performance trade-off. Together, PatternEval and PatternRL provide an evaluation-and-training framework for aligning user-visible response patterns across hybrid-thinking interfaces.
Problem

Research questions and friction points this paper is trying to address.

response-pattern alignment
hybrid-thinking MLLMs
response behavior
multimodal reasoning
failure modes
Innovation

Methods, ideas, or system contributions that make the work stand out.

response-pattern alignment
PatternEval
PatternRL
hybrid-thinking MLLMs
multimodal reasoning failures
X
Xinming Wang
Institute of Automation, Chinese Academy of Sciences
Weinong Wang
Weinong Wang
Xian Jiaotong University
LLM/VLLM/RL
H
Hongming Yang
Large Language Model Department, Tencent
Y
Yansong Lin
University of Electronic Science and Technology of China
Z
Zheng Ruan
Large Language Model Department, Tencent
Shangpin Peng
Shangpin Peng
Harbin Institute of Technology, Shenzhen
Artificial IntelligenceLLMPreference Optimization
Q
Qiming Peng
Large Language Model Department, Tencent
Nan Qiao
Nan Qiao
Amazon
Semantic SegmentationRepresentation LearningActive Learning
F
Fengyuan Lu
Large Language Model Department, Tencent
G
Guoqing Ma
Large Language Model Department, Tencent
M
Marito Li
Large Language Model Department, Tencent
S
Songyang Zhang
Large Language Model Department, Tencent
S
Saiyong Yang
Large Language Model Department, Tencent
Han Hu
Han Hu
Distinguished Scientist, Tencent Hunyuan
Computer VisionDeep LearningMachine Learning
Yonglong Tian
Yonglong Tian
Research Scientist, OpenAI
Artificial IntelligenceDeep LearningMachine LearningComputer Vision
Xu-Yao Zhang
Xu-Yao Zhang
Institute of Automation, Chinese Academy of Sciences
Pattern RecognitionMachine LearningOCR