Hierarchical Wasserstein Merging for Multi-Domain Multi-Task Learning: From Specialists to a Generalist

📅 2026-09-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决多领域多任务学习中的分布偏移干扰问题,提出了一种基于表示层次的Hierarchical Wasserstein Merging方法,通过构建Wasserstein重心来聚合专家模型或训练通用模型。
📝 Abstract
Multi-domain multi-task learning (MD-MTL) aims to build a single generalist model that performs well across heterogeneous domains and tasks. However, joint training often suffers from interference under distribution shifts. Existing model merging methods mostly operate on model parameters while overlooking the geometric structure of latent representation distributions across domains and tasks. To address these limitations, we propose Hierarchical Wasserstein Merging (HWM), a representation-level framework that models each domain-task specialist as a distribution of hidden representations on a shared support. HWM constructs task-level and global Wasserstein barycenters to capture within-task domain variation and cross-task structure, enabling either training-free specialist aggregation by Wasserstein-derived weights or training-based generalist learning through a hybrid Wasserstein alignment loss. Experiments on four NLP tasks across four domains per task show that HWM achieves superior effectiveness and generalization capability in MD-MTL settings.
Problem

Research questions and friction points this paper is trying to address.

Multi-domain Multi-task Learning
Distribution Shifts
Latent Representation Distributions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical Wasserstein Merging
Representation-level Framework
Wasserstein Barycenters
Multi-Domain Multi-Task Learning
🔎 Similar Papers
No similar papers found.