ExaServe: Large-Scale LLM Serving on Exascale HPC Systems

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
ExaServe解决了在超级计算机上部署大规模语言模型(LLM)的工程挑战,通过一个可pip安装的框架将YAML规范转换为可重复的大规模LLM服务部署。
📝 Abstract
Cloud-native LLM serving frameworks have made deployment routine in data centers, yet deploying them on leadership-class supercomputers remains an engineering challenge requiring scheduler integration, MPI launch, accelerator selection, node-local weight staging, and platform-specific patches. We present \textit{ExaServe}, a pip-installable framework that transforms a declarative YAML specification into a reproducible large-scale LLM serving deployment. Using ExaServe, we deploy LLM serving on ALCF Aurora from 1 to 256 nodes (3072 vLLM replicas). Non-streaming inference scales nearly linearly to 256 nodes, reaching 27.1\,k requests/s (3.8\,M tokens/s). Token streaming scales differently: a centralized proxy plateaus at $\sim$4.7\,k requests/s despite the model servers remaining within the service-level objective. We also identify an \emph{O}(\emph{N}\textsuperscript{2}) Ray Serve control-plane bottleneck that increases cluster bring-up to $\sim$30 minutes at 256 nodes. ExaServe provides a practical, reproducible deployment path while exposing key barriers to future exascale LLM serving.
Problem

Research questions and friction points this paper is trying to address.

LLM serving
supercomputers
engineering challenge
scheduler integration
MPI launch
Innovation

Methods, ideas, or system contributions that make the work stand out.

Exascale HPC Systems
Large-Scale LLM Serving
YAML Specification
Non-streaming Inference Scaling
Control-Plane Bottleneck