A.X K2 Technical Report

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文介绍了A.X K2,一个688B参数的语言模型,通过Sparse Gated Attention和Gated Norm技术提高长文本处理效率与质量,超越前代模型。
📝 Abstract
We introduce A.X K2, a 688B-parameter Mixture-of-Experts (MoE) language model trained from scratch as a high-performance foundation for \emph{agentic} applications. Trained on approximately 8.5T tokens---fewer than its predecessor, A.X K1---on a smaller but higher-quality mixture with substantially expanded agentic and software-engineering data, it nonetheless improves over A.X K1 across the board, by over 30 percentage points on some benchmarks, reflecting large gains in token efficiency. To support long contexts efficiently, we introduce Sparse Gated Attention (SGA), which combines sparse attention with gated attention, and adopt Gated Norm (GN) to stabilize large-scale training. SGA is trained natively at 128K through a \emph{sparse} indexer warmup that optimizes the indexer against its own sparse top-$k$ selection rather than the dense attention distribution, making adaptation markedly cheaper: each query reads only 2,048 positions, yet long-context quality is unchanged and A.X K2 scores 94.6 on RULER out to 256K. The outlier suppression of GN in turn keeps 4-bit NVFP4 serving within one point of FP8 accuracy. A simple yet effective Think-Fusion recipe further lets users switch between thinking and non-thinking modes within a single unified model. Extensive evaluations show that A.X K2 performs competitively against strong open-weight baselines, matching or exceeding them on math and Korean-language benchmarks.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
agentic applications
long contexts
token efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Sparse Gated Attention
Gated Norm
Think-Fusion
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Cheolseung Baek
SK Telecom
D
Dhammiko Arya
SK Telecom
Eunki Kim
Eunki Kim
KAIST
AIMulti-modalNLP
G
Gun Song
SK Telecom
G
Gyoungeun Han
SK Telecom
H
Hyunho Yang
SK Telecom
Hyunjun Eun
Hyunjun Eun
SK Telecom
Deep LearningComputer Vision
J
Jin Kim
SK Telecom
J
Junyoung Park
SK Telecom
J
Juyun Wee
SK Telecom
Minki Hong
Minki Hong
Korea Advanced Institute of Science and Technology
Human-Computer InteractionComputational Interaction
Minkyung Park
Minkyung Park
SK Telecom
Minsang Kim
Minsang Kim
Korea university
Machine LearningDeep LearningNLPLLMFoundation Models
Minsoo Kang
Minsoo Kang
SK Telecom
Machine LearningComputer Vision
S
SaeRom Kim
SK Telecom
S
Sangjin Kim
SK Telecom
Sangyeol Lee
Sangyeol Lee
Seoul National University
Financial Time SeriesRisk managementChange point analysis & Statististical Process ControlPredictive Analytics
S
Seojin Lee
SK Telecom
S
Seokhwan Jo
SK Telecom
S
Seokyoung Hong
SK Telecom
Seongho Choi
Seongho Choi
Ph. D Student at Seoul National University
Video QA
S
Seonghye Cho
SK Telecom
S
Seongmin Ok
SK Telecom
S
Sereimony Sek
SK Telecom
S
Seungmo Cho
SK Telecom