NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过引入NCP-ArchPreview模型,采用超越传统的下个概念预测方法来改进语言模型的预训练过程,以解决现有模型在概念层面理解不足的问题。
📝 Abstract
We introduce NCP-ArchPreview, a latent-space language model that pushes autoregressive pretraining beyond standard next-token prediction (NTP). Alongside NTP, the model learns through Next Concept Prediction (NCP) to predict discrete concepts that span multiple tokens, introducing an explicit and more challenging concept-level objective while preserving standard token-level autoregressive generation. NCP-ArchPreview builds a latent space by constructing a product-quantized concept vocabulary directly from its hidden states, and subsequently learns to predict future concepts via a dedicated Concept Module. These predicted concepts are then fed back to the token level to guide subsequent generation, with NTP and NCP trained jointly end-to-end. We scale this architecture to 8.9B parameters and train it on 5.73T tokens from the Dolma-3 dataset, marking the largest demonstration of a latent-space language model to date. Remarkably, by consuming only 51.3% of the total training tokens, NCP-ArchPreview achieves the final pretraining loss of OLMo-3-7B. Following full pretraining, it outperforms OLMo-3-7B by 2.45 points on the downstream macro-average, including a notable 5.99-point gain on GSM8K. Controlled experiments isolate a clear progression of performance gains stemming from both the latent architecture and the NCP objective. Furthermore, utilizing only 85% of the standard computation, NCP-ArchPreview approaches the training loss of a strictly parameter-aligned 8.9B baseline. The learned latent space remains highly valuable after the pretraining stage: updating just the 17M-parameter VQ module yields a novel, lightweight interface for domain adaptation, while a simple injection of concept representations into a DFlash2 drafter improves the mean accepted length by 4.17% with negligible overhead.
Problem

Research questions and friction points this paper is trying to address.

Next Concept Prediction
Latent Space Language Model
Autoregressive Pretraining
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent Space
Next Concept Prediction
Concept Module
Autoregressive Pretraining
Domain Adaptation
🔎 Similar Papers
T
The Intern-NCP Team
Shanghai AI Lab; LUMIA Lab, Shanghai Jiao Tong University
Jiaqi Cao
Jiaqi Cao
Shanghai Jiao Tong University
Natural Language ProcessingLong-term Memory
C
Chiyu Chen
S
Shuang Cheng
X
Xu Cheng
B
Beiya Dai
Y
Yufan Feng
K
Kewen Ge
R
Ruijun Ge
J
Jiayi Huang
Y
Yang Jiao
Dahua Lin
Dahua Lin
The Chinese University of Hong Kong
computer visionmachine learningprobabilistic inferencebayesian nonparametrics
Z
Zhouhan Lin
Y
Yifan Liu
Y
Yuliang Liu
B
Biqing Qi
M
Mowen Ruan
J
Junzhe Shen
Yunchong Song
Yunchong Song
Ph.D. student, Shanghai Jiao Tong University
Machine Learning
H
Hao Sun
Z
Zhongbo Tian
Y
Yixuan Wang
Rubin Wei
Rubin Wei
Shanghai Jiao Tong University
LLMMemory-Augmented LLM
J
Jiaxin Xiong