Towards a Densing Law for User Representation Learning at Billion-Scale Capacity

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出用户行为稠密定律,通过量化分析解决十亿级规模下原始数据扩展瓶颈和token化配置问题,开发了自适应变长token化方法ALGN以优化容量分配。
📝 Abstract
User representation learning in real-world industrial scenarios is commonly scaled by increasing user amount, behavioral sequence length and model size. However, existing methods face two challenges: (i) Bottleneck for raw data scaling at billion-scale capacity, as performance exhibit diminishing performance gains with larger-scale raw text user behavioral input, which can be mitigated by tokenization. (ii) Lack of quantitative analysis of how tokenization configurations should scale with data size. In this report, we propose User Behavioral Densing Law for characterizing the quantitative relationship between data scale and the minimum sufficient tokenization capacity. Firstly, we conduct a pilot study on raw & tokenized scaling comparison on billion-scale Alipay dataset, revealing the raw data scaling bottleneck and the sustained gains enabled by tokenization. To derive the scaling pattern governing the minimum sufficient tokenization configuration at different data scales, theoretical analysis and systematic experiments are employed to summarize the quantitative scaling pattern. We find an approximately linear relationship between the logarithms of minimum sufficient tokenization capacity and input data size measured by tokens, and the scaling slope varies systematically with the tokenization method and data source, reflecting differences in representation-space redundancy and intra-source uniqueness. Guided by the proposed law, we further develop ALGN, an adaptive variable-length tokenization method that improves capacity allocation. Extensive experiments across diverse data sources, tokenization methods, and downstream tasks demonstrate the generalizability and reliability of the User Behavioral Densing Law, providing practical guidance for tokenization configuration selection in large-scale user representation learning. Moreover, ALGN outperforms existing baselines.
Problem

Research questions and friction points this paper is trying to address.

User Representation Learning
Billion-Scale Capacity
Tokenization
Data Scaling
Behavioral Sequence
Innovation

Methods, ideas, or system contributions that make the work stand out.

User Behavioral Densing Law
tokenization configuration
adaptive variable-length tokenization
capacity allocation
large-scale user representation learning
🔎 Similar Papers
B
Bin Dou
DeepFind Team, Ant Group
Junru Zhang
Junru Zhang
Zhejiang University
LLMsData MiningTime SeriesTime Series ClassificationDomain Generalization
Z
Zhaoyi Yuan
DeepFind Team, Ant Group; Zhejiang University
W
Wuliang Huang
DeepFind Team, Ant Group
L
Letian Gong
DeepFind Team, Ant Group
B
Baokun Wang
DeepFind Team, Ant Group
H
Huan Li
Zhejiang University
Y
Yu Cheng
DeepFind Team, Ant Group
W
Weiqiang Wang
DeepFind Team, Ant Group