TailSieve: Partial-Rollout-Guided Tail Routing for LLM Rollouts

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决大规模语言模型部署中长尾生成导致的整体延迟问题,提出TailSieve框架,通过部分展开指导和自适应资源分配优化路由策略。
📝 Abstract
Large-scale rollouts have become a core component of modern LLM systems, spanning reinforcement learning (RL) post-training, on-policy distillation (OPD), and sampling-heavy evaluation pipelines. Unlike online serving, which is typically optimized for request-level latency and throughput, a small number of long-tail generations can dominate the end-to-end makespan of an entire rollout step. In practice, rollout requests are often routed uniformly across replicas, which can place extremely long generations inside high-concurrency decoding batches. To address this, we present TailSieve, a partial-rollout-guided framework that jointly controls tail routing and replica allocation for LLM rollouts. In an idealized setting with known completion lengths, we show that makespan-optimal routing in the long-tail regime combines tail isolation with load balancing, and that a simple top-k policy closely approximates this offline optimum. Leveraging the observation that long-tail prompts tend to remain long-tailed across policy updates, TailSieve uses partial rollouts as a training-free signal for identifying candidate tail groups. A hierarchical controller then jointly adapts the number of isolated groups and the replica split between the tail and bulk pools using collected response-work history and a measured concurrency-throughput model. TailSieve achieves up to 1.67x routing-only speedup over uniform group routing. The resulting low-concurrency tail pool further enables route-specialized speculative decoding with MTP or DFlash, achieving up to 2.59x speedup over uniform routing. Selected prompts are regenerated under the current policy, preserving on-policy generation and avoiding additional routing-induced length bias in steady state.
Problem

Research questions and friction points this paper is trying to address.

large-scale rollouts
long-tail generations
replica allocation
makespan
tail routing
Innovation

Methods, ideas, or system contributions that make the work stand out.

partial-rollout-guided
tail routing
replica allocation
load balancing
speculative decoding
🔎 Similar Papers
No similar papers found.
Tianqi Xu
Tianqi Xu
Tokyo Institute of Technology
CloudBurst BufferDistributed File SystemUser-Level File SystemHPC
L
Lu Lv
Qwen Business Unit of Alibaba
Haoyang Huang
Haoyang Huang
JD Explore Academy (present) | StepFun | Microsoft Research
Multimodal & Multilingual Foundation Model
Wenjie Huang
Wenjie Huang
Shanghai Jiao Tong University
点云压缩视频压缩图像压缩
Z
Zhanming Shen
Zhejiang University
Y
Yuhao Shen
Zhejiang University
B
Baolin Zhang
Qwen Business Unit of Alibaba
X
Xinyi Hu
Qwen Business Unit of Alibaba
S
Shuang Ge
Qwen Business Unit of Alibaba
J
Jun Dai
Qwen Business Unit of Alibaba
Tianyu Liu
Tianyu Liu
Hongkong University of Science and Technology
Suorong Yang
Suorong Yang
Nanjing University
Computer VisionDeep LearningMultimodal Learning
Zhikai Li
Zhikai Li
Assistant Professor, Institute of Automation, Chinese Academy of Sciences
Efficient Deep LearningModel CompressionComputer VisionHW-SW Co-Design
Y
Ye Bai
Qwen Business Unit of Alibaba
J
Jun Zhang
Qwen Business Unit of Alibaba
L
Lei Chen
Qwen Business Unit of Alibaba
Y
Yue Li
Qwen Business Unit of Alibaba
M
Mingchen Wan
Qwen Business Unit of Alibaba