MCPO: Modality-Contrastive Preference Optimization for Multimodal Chain-of-Thought Compression

๐Ÿ“… 2026-09-04
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡ๆๅ‡บMCPOๆ–นๆณ•๏ผŒ้€š่ฟ‡ไธค้˜ถๆฎตๅŽ‹็ผฉๅ’Œๅฏน้ฝไผ˜ๅŒ–ๅคšๆจกๆ€้•ฟ้“พๆŽจ็†๏ผŒๅ‡ๅฐ‘่ฎก็ฎ—ๆˆๆœฌๅ’Œๅ†—ไฝ™๏ผŒๆ้ซ˜ๆŽจ็†ๆ•ˆ็އใ€‚
๐Ÿ“ Abstract
Recently, multimodal large-scale reasoning models have demonstrated remarkable capabilities in solving complex tasks through long Chains-of-Thought (M-CoT). However, excessively long reasoning trajectories incur substantial computational costs and significant KV-cache pressure. Existing CoT compression and alignment paradigms mainly rely on static rules or single-dimensional preferences, lacking fine-grained cross-modal constraints; as a result, they are prone to inducing visual laziness and hallucinatory reasoning. To address these issues, we propose Modality-Contrastive Preference Optimization (MCPO), a highly sample-efficient two-stage length-compression method that requires fewer than 900 training samples. In the compression stage, we introduce a step-level Normalized Cross-Modal Mutual Information (NCMI) pruning algorithm, which automatically identifies and removes visual-independent reasoning steps by comparing the reasoning discrepancies between with-image and no-image contexts. This significantly reduces redundancy and hallucinatory content in the reasoning chains. In the alignment stage, the model first undergoes supervised fine-tuning to achieve domain-adaptive initialization, followed by optimization using an asymmetric multimodal length-controlled preference loss. This objective adopts a highly nonlinear odds-ratio formulation that provides steep gradients in the with-image context to reinforce length constraints for preferred trajectories, while applying a scaled, flat-gradient linear difference in the no-image context to maintain modality consistency, thereby achieving stable cross-modal preference alignment. Extensive experiments on mainstream base models such as Qwen3-VL-Thinking show that our method can reduce CoT length by up to 69.5% and achieve up to 3.34x end-to-end inference speedup while preserving original accuracy.
Problem

Research questions and friction points this paper is trying to address.

multimodal reasoning
computational costs
KV-cache pressure
cross-modal constraints
visual laziness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Modality-Contrastive Preference Optimization (MCPO)
Normalized Cross-Modal Mutual Information (NCMI) pruning
asymmetric multimodal length-controlled preference loss
Chain-of-Thought (CoT) compression
multimodal reasoning
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
G
Guangheng Yang
Tsinghua University, Huawei Technologies Ltd.
Z
Zhenliang Ni
Huawei Technologies Ltd.
Z
Zhenkai Wu
Huawei Technologies Ltd.
Han Shu
Han Shu
Huawei Noah's Ark Lab
Juan Feng
Juan Feng
Tsinghua University
Wenming Yang
Wenming Yang
Tsinghua University
Computer VisionImage Processing
J
Jie Hu
Huawei Technologies Ltd.