End-to-End Latency-Minimizing and Load-Balanced Request Scheduling for Edge LLM Inference in Agentic AI Services

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文研究了边缘LLM推理的请求调度框架,通过LYREO方法最小化端到端延迟并平衡负载分布,解决了动态变化下的请求调度难题。
📝 Abstract
Large language model (LLM)-powered agentic AI services increasingly demand low-latency inference, motivating the deployment of LLMs across distributed edge servers. However, heterogeneous communication and computing capabilities, together with dynamically evolving inference states, make the edge server selection for each incoming request time-varying and tightly coupled across slots. In this paper, we investigate an online request scheduling framework for edge LLM inference that jointly minimizes long-term average end-to-end latency and regulates workload distribution across heterogeneous edge servers. Two main challenges arise in this context. First, conventional latency models cannot accurately capture the fine-grained dynamics of multi-stage LLM execution. Second, the latency consequence of a scheduling decision is observed only after request completion, making immediate decision evaluation difficult. To address these challenges, we develop a cross-slot inference model that captures transmission, prefill, iteration-level decoding, and key-value (KV) cache evolution for each diverse request, and characterize server workload through a KV cache memory-time consumption metric. We propose the LYREO approach that transforms the long-term load-balancing constraint via Lyapunov optimization and employs reward redistribution with sequencebased return prediction to convert delayed outcomes into timely learning signals for earlier decisions. Simulations under various configurations demonstrate that LYREO consistently achieves lower latency and more balanced load distribution than representative learning-based and heuristic baseline schemes.
Problem

Research questions and friction points this paper is trying to address.

end-to-end latency
load balancing
edge LLM inference
heterogeneous edge servers
request scheduling
Innovation

Methods, ideas, or system contributions that make the work stand out.

LYREO
Lyapunov optimization
reward redistribution
cross-slot inference model
KV cache evolution
Z
Zhen Li
Department of Electrical and Computer Engineering, Concordia University, Montreal, QC, H3G 1M8, Canada
Jun Cai
Jun Cai
Concordia University
Wireless Communication Networks
H
Haoran Gao
Department of Electrical and Computer Engineering, Concordia University, Montreal, QC, H3G 1M8, Canada
An Li
An Li
Futurewei Technologies, Inc.
Optical CommunicationsFiber Sensors
T
Tan Li
Department of Computer Science, Hang Seng University of Hong Kong, Hong Kong SAR