MAPS: Memory-Aware Predictive Scheduling Framework for Large Language Model Serving

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决云服务基础设施中的负载不平衡问题,提出MAPS框架,通过预测输出长度和分级调度策略优化大语言模型的服务性能。
📝 Abstract
The surge of large language model (LLM) applications on personal devices imposes massive, bursty workloads on cloud serving infrastructure. While prefill-decode disaggregation improves throughput and scalability, memory-bound decode instances often suffer from persistent load imbalance, as output lengths are unknown when requests arrive at the cloud. To address this, we propose MAPS, a Memory-Aware Predictive Scheduling framework tailored for disaggregated LLM serving. MAPS performs device-assisted speculative output length prediction overlapped with cloud-side prefilling, incurring negligible latency overhead. To handle generation uncertainty, MAPS applies uncertainty-aware calibration to derive output-length upper bounds with target coverage, enabling safe scheduling decisions. Building on these bounds, MAPS employs a hierarchical global-local scheduling strategy to mitigate inter-decoder queue buildup and intra-decoder head-of-line blocking. Extensive experiments on two real-world workloads and two LLMs show that MAPS significantly outperforms three state-of-the-art systems, reducing average end-to-end latency by 42.6 and tail latency by up to 84.8.
Problem

Research questions and friction points this paper is trying to address.

large language model
cloud serving infrastructure
load imbalance
output length prediction
memory-aware scheduling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Memory-Aware Predictive Scheduling
speculative output length prediction
uncertainty-aware calibration
hierarchical global-local scheduling
🔎 Similar Papers
Tiancheng Zhang
Tiancheng Zhang
Northeastern University, China
user profiledeep learningmachine learning,intelligent education
Y
Yulin Chen
The International Joint Institute of Tianjin University, Tianjin University, Fuzhou, China
Yunfeng Zhao
Yunfeng Zhao
Tianjin University
Edge computing
S
Shaoyuan Huang
College of Intelligence and Computing, Tianjin University, Tianjin, China
C
Cheng Zhang
Faculty of Digital Economics and Managements, Tianjin University of Finance and Economics, Tianjin, China
Xiaofei Wang
Xiaofei Wang
Professor, Tianjin University; Chief Scientist, PPIO
edge computingedge intelligencemobile social networksnetwork big datasmart city and IoT