Can Released LLM Vocabularies Support Token-Level Estimation of Hidden Corpora?

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of estimating the true proportion of arbitrary tokens in the typically undisclosed pretraining corpora of large language models. The authors propose a novel transfer-based estimation method leveraging the vocabulary of Byte Pair Encoding (BPE) tokenizers, enabling fine-grained inference of target token frequencies in hidden corpora for the first time. The core innovation lies in the introduction of quantile-guided density estimation (QGDE) combined with a local density weighting strategy, which substantially enhances estimation accuracy. Experimental results demonstrate that the approach achieves remarkably low average relative errors—3.00% at the token level and 3.08% after category-level aggregation—across both controlled settings and real-world scenarios, such as with the SmolLM tokenizer.
📝 Abstract
Pretraining corpus composition shapes LLM capabilities, but it often remains hidden even when model weights are released. Prior work has inferred corpus mixtures or traced specific token groups from released tokenizer vocabularies; in contrast, we estimate corpus ratios for arbitrary target tokens. We first show that BPE tokenizers trained on different corpora share stable token ID--ratio distributions, motivating distribution transfer from known corpora to a target tokenizer trained on hidden corpora. We then propose Quantile-Guided Density Estimation (QGDE), which approximates this distribution with multiple quantile trends and uses local density weighting to produce token-level estimates. In controlled settings and a realistic setting using the released SmolLM tokenizer, QGDE achieves mean relative errors as low as 3.00% for token-level estimation and 3.08% after aggregation into category-level mixtures. These results suggest that released tokenizer vocabularies provide a useful signal for fine-grained corpus estimation beyond coarse composition inference.
Problem

Research questions and friction points this paper is trying to address.

pretraining corpus
tokenizer vocabulary
token-level estimation
hidden corpora
corpus composition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Quantile-Guided Density Estimation
tokenizer vocabulary
corpus composition estimation
BPE tokenizer
token-level inference
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Q
Qingjie Zhang
Tsinghua University; Qwen Team, Alibaba Group
X
Xingzhang Ren
Qwen Team, Alibaba Group
Z
Zixuan Chen
Tsinghua University
Jinfeng Li
Jinfeng Li
Alibaba Group
AI SecurityTruthworthy AI
Y
YueFeng Chen
Alibaba Group
Yitong Yang
Yitong Yang
Shanghai University of Finance and Economics
H
Hui Xue
Alibaba Group
D
Dayiheng Liu
Qwen Team, Alibaba Group
Han Qiu
Han Qiu
NTU