Kimi K3: Open Frontier Intelligence

πŸ“… 2026-07-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses performance and efficiency bottlenecks of large language models in long-context, multimodal, and complex reasoning tasks by introducing a 2.8 trillion-parameter sparsely activated Mixture-of-Experts model with native vision capabilities and support for million-token contexts. Key innovations include Kimi Delta Attention and Attention Residuals to enhance information flow, Stable LatentMoE for robust and efficient expert routing, and a co-designed algorithm-system reinforcement learning framework featuring persistent rollouts and sandboxed states. The model achieves state-of-the-art performance across long-context encoding, agent-based tasks, knowledge-intensive question answering, reasoning, and vision benchmarks, demonstrating a 2.5Γ— improvement in scaling efficiency over its predecessor. It surpasses existing open-source models and most closed-source counterparts, with full model weights publicly released.
πŸ“ Abstract
We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model depth. Together with Stable LatentMoE, which effectively activates 16 of 896 routed experts per token, and refined training and data recipes, these advances yield an approximately 2.5x improvement in overall scaling efficiency over Kimi K2. Post-training highlights reinforcement learning across general, agentic, and coding domains and multiple reasoning-effort levels, enabling compositional generalization and robust long-horizon execution. At 2.8T scale, Kimi K3 is supported by infrastructure advances in multiple areas: algorithm-system co-design for KDA, perfectly balanced expert-parallel training with efficient memory management, million-token agentic RL with persistent rollout and sandbox states, and deployment innovations. Extensive evaluations show that Kimi K3 achieves frontier-level performance across long-horizon coding, agentic, knowledge, reasoning, and vision tasks. While its overall performance still trails the most powerful proprietary models, namely Claude Fable 5 and GPT-5.6 Sol, Kimi K3 consistently outperforms other open and proprietary models evaluated in our suite. We release the full Kimi K3 model weights to facilitate future research and accelerate the broader deployment and adoption of frontier intelligence.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
long-context modeling
frontier intelligence
vision-language models
scaling efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Delta Attention
Stable LatentMoE
million-token context
agentic reinforcement learning
πŸ”Ž Similar Papers
No similar papers found.
K
Kimi Team
Moonshot AI (ζœˆδΉ‹ζš—ι’)