MiMo-V2-Flash Technical Report
This work proposes a 309B-parameter sparse mixture-of-experts (MoE) language model with only 15B activated parameters per token, designed to enhance reasoning speed, capability, and agent-task performance while reducing computational costs. The architecture integrates sliding-window and global attention mechanisms and introduces a multi-token prediction (MTP) framework alongside a multi-teacher online policy distillation (MOPD) approach to enable efficient training and speculative decoding. Despite using merely one-half to one-third of the activated parameters compared to leading open-source models of similar scale, the proposed model achieves comparable or superior performance, accelerates inference by up to 2.6×, and supports context lengths of up to 3.6 million tokens.