Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过比较三种融合范式(Merge、Mix RL、MOPD)来解决多领域能力整合问题,指导选择合适的方法以提升大型语言模型特定能力。
📝 Abstract
Reinforcement learning with verifiable rewards (RLVR) improves specific capabilities of large language models, but covering multiple capabilities often involves training separate domain experts and subsequently consolidating them. We organize three fusion paradigms by the artefacts they reuse: Merge combines expert task vectors, Mix RL pools their datasets, and multi-teacher on-policy distillation (MOPD) uses both. Because they have largely been studied in isolation, how they compare and how to choose among them remain unclear. We compare all three using shared experts and data across model scales and a multi-domain benchmark suite. Although their average performance differs by at most 1.4 points, the gap reaches 8.6 points on a single benchmark, with domain-level variation tracking cross-domain relations visible in task-vector geometry. Training dynamics expose distinct constraints: Mix RL depends on domain mixture proportions, MOPD remains bounded by its teachers, and Merge compresses all expert updates into one. All three improve single-sample accuracy without measurable gains in solution coverage or losses in held-out capabilities. These results yield a practical guideline: use Merge when experts already exist and cheap fusion is paramount; Mix RL when training a unified model without experts, with domain proportions adjusted for cross-domain transfer; and MOPD when preserving domain-specific gains matters more than surpassing teachers or minimizing end-to-end cost.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Verifiable Rewards
Fusion Paradigms
Domain Experts
Consolidation
Innovation

Methods, ideas, or system contributions that make the work stand out.

RLVR
fusion paradigms
cross-domain
task-vector geometry
training dynamics
🔎 Similar Papers
Siye Wu
Siye Wu
Fudan University
K
Kai Yang
LLM Department, Tencent
Y
Yuchen Cai
LLM Department, Tencent
X
Xin Xu
LLM Department, Tencent
P
Peng-Yuan Wang
LLM Department, Tencent
Jiaxuan Wang
Jiaxuan Wang
GE HealthCare
Model interpretabilityMachine learning for healthcareOut of distribution generalization
J
Jiashun Liu
LLM Department, Tencent
Jiafei Lyu
Jiafei Lyu
PhD of Control Science and Engineering, Tsinghua University
deep reinforcement learning
Y
Yangkun Chen
LLM Department, Tencent
S
Saiyong Yang
LLM Department, Tencent
Y
Yanghua Xiao
Fudan University