Accelerating Unified Multimodal Models with Core-Expansion Routing and Unified Computation Scheduling

📅 2026-08-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过提出CE-Router和统一计算调度方法,解决了多模态模型中冗余计算的问题,提高了推理速度同时保持了性能。
📝 Abstract
Unified multimodal models jointly support understanding and generation, but incur substantial redundant computation across tokens, layers, and generation timesteps. Through token-importance probing, we identify an asymmetric core-expansion structure: understanding exhibits a stable importance component, while generation largely shares this component but requires progress-dependent corrections. We therefore propose CE-Router, which uses a task-shared core scorer and progress-conditioned generation expansions, optimized through generation decomposition and cross-task core alignment. At inference, CE-Router compacts token computation and supplies a learned routing signal to Unified Computation Scheduling, which coordinates layer skipping, FFN pruning, diffusion-head cache reuse, and denoising-step early exit. Experiments on two representative UMM architectures demonstrate consistent quality--efficiency improvements across both tasks, retaining 98.03\% of dense understanding performance with a 1.93$\times$ end-to-end inference speedup.
Problem

Research questions and friction points this paper is trying to address.

Unified multimodal models
redundant computation
tokens
layers
generation timesteps
Innovation

Methods, ideas, or system contributions that make the work stand out.

CE-Router
Unified Computation Scheduling
Core-Expansion Routing
Token-Importance Probing