SequenceO1: End-to-End Ultra-Long (100K) Sequence Modeling in Recommendation with Low-Rank Caching

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究解决了长序列推荐系统中的效率问题,通过SequenceO1框架使用低秩缓存和压缩技术处理高达100K的用户行为序列。
📝 Abstract
Modeling long-term user behavior is central to sequential recommendation and billion-scale industrial recommender systems, yet production ranking models operate under strict latency, memory, communication, and training-throughput constraints. At the 100K scale, the challenge extends beyond attention complexity: raw sequence features must be stored, transferred, and repeatedly processed during training and online serving. Existing approaches based on history truncation, multi-stage behavior retrieval, compressed lifelong histories, or train-short/infer-long extrapolation either weaken end-to-end optimization or retain substantial length-dependent cost. We present SequenceO1, an end-to-end framework for ultra-long user behavior sequence modeling, deployed at full traffic on Douyin with histories of up to 100K interactions. SequenceO1 follows a compress-then-reason design. Its Sketch Attention (SA) uses learnable prototypes and prototype-wise normalization to compress the raw history into a fixed-size, target-agnostic user representation. Target-conditioned Stacked Target-to-History Cross Attention (STCA) then models complementary time scales: a recent 10K suffix for short-term interests and the compact sketch for long-term preferences. To make training and inference practical, SequenceO1 combines low-rank user representation caching, multi-request user-level batching, pipeline lift, and a fused FlashSA kernel to amortize feature storage, communication, and computation across targets, training instances, and consecutive requests. Production experiments show consistent offline and online gains, while the compact cached sketch retains most of the benefit of directly scaling end-to-end sequence ranking to 100K. These results provide a practical model-system approach to efficient attention, sequence compression, and scalable long-sequence and long-context recommendation systems.
Problem

Research questions and friction points this paper is trying to address.

long-term user behavior
recommendation systems
latency constraints
sequence modeling
attention complexity
Innovation

Methods, ideas, or system contributions that make the work stand out.

End-to-End
Ultra-Long Sequence Modeling
Low-Rank Caching
Sketch Attention
Stacked Target-to-History Cross Attention
🔎 Similar Papers
L
Lin Guan
ByteDance, Beijing, China
Jia-Qi Yang
Jia-Qi Yang
ByteDance
machine learningdata miningrecommender systems
Z
Zhishan Zhao
ByteDance, Beijing, China
Jiaqi Huang
Jiaqi Huang
University of Central Missouri
CybersecurityIoV
Hangyu Wang
Hangyu Wang
Shanghai Jiao Tong University
Information RetrievalRecommender System
L
Longbin Li
ByteDance, Beijing, China
Beichuan Zhang
Beichuan Zhang
Professor of Computer Science, the University of Arizona
Computer Networks
H
Haonan Jiang
ByteDance, Shanghai, China
J
Jinan Ni
ByteDance, Shanghai, China
X
Xiangyu Fan
ByteDance, Shanghai, China
X
Xiaowen Li
ByteDance, Beijing, China
Z
Ziyao Ren
ByteDance, Beijing, China
Y
Yuhang Qi
ByteDance, Hangzhou, Zhejiang, China
X
Xiaolong Zhu
ByteDance, Beijing, China
X
Xuanyuan Luo
ByteDance, Hangzhou, Zhejiang, China
Q
Qiwei Chen
ByteDance, Shanghai, China
Y
Yi Cheng
ByteDance, Beijing, China
L
Lele Yu
ByteDance, San Jose, CA, USA