MoTE: Mixture of Task Experts for Multi-Task Video Understanding

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决多任务视频理解中任务行为纠缠问题,提出MoTE架构,通过将大型语言模型前馈网络转换为任务特定专家来提高计算效率和准确性。
📝 Abstract
Procedural video-language models must solve heterogeneous tasks from the same visual evidence, including action recognition, forecasting, and procedure prediction. Dense transformer decoders share the same feed-forward networks across tasks, which can entangle task behavior and make controlled capability expansion difficult. Sparse Mixture-of-Experts (MoE) decoders provide conditional computation, but token-level learned routing is not naturally aligned with task-level procedural objectives. We propose MoTE (Mixture of Task Experts), a decoder architecture that converts large language model feed-forward networks into task-specific experts while keeping the multimodal backbone shared. Each example follows one sample-level task route, so active task-expert computation remains independent of the number of stored task experts. We instantiate this design as VideoLLM-MoTE and evaluate it on five COIN benchmarks using explicit task routes. The five-expert model activates ~2B LLM parameters per sample and achieves higher average top-1 accuracy than recent VideoLLM baselines. Under the same expert topology, it improves over dense all-expert activation and learned sparse-routing controls. These results show that task-structured routing provides an interpretable and compute-efficient decoder alternative for multi-task video-language learning.
Problem

Research questions and friction points this paper is trying to address.

Multi-Task Video Understanding
Dense Transformer Decoders
Mixture-of-Experts (MoE)
Conditional Computation
Task-Specific Experts
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture of Task Experts
task-specific experts
multimodal backbone
task-structured routing
multi-task video-language learning
🔎 Similar Papers
2024-02-20International Conference on Machine LearningCitations: 30
M
Muhammad Asad Ali
Department of Computer Science, University of Kaiserslautern-Landau (RPTU), Kaiserslautern, Germany; Augmented Vision Group, German Research Center for Artificial Intelligence (DFKI), Kaiserslautern, Germany
U
Umar Khan
Department of Computer Science, University of Kaiserslautern-Landau (RPTU), Kaiserslautern, Germany
N
Nadia Robertini
Augmented Vision Group, German Research Center for Artificial Intelligence (DFKI), Kaiserslautern, Germany
Didier Stricker
Didier Stricker
Professor for Computer Science, University Kaiserslautern
augmented realitycomputer visionimage processingbody sensor networkshci