Global Simulation-Guided Dynamic Operator Scheduling for Efficient Multi-Tenant Model Serving

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses resource fragmentation in container scheduling and the challenges of operator scheduling dependencies and memory safety in multi-tenant model serving. We propose SliceScheduler, a system that leverages global mapping graph abstraction and incremental simulation to perceive cluster states in real time and enable dynamic operator-level scheduling. This approach achieves practical fine-grained GPU resource multiplexing while maintaining Service Level Agreement (SLA) guarantees for the first time. Experimental results demonstrate that SliceScheduler increases token throughput by 1.10× to 2.29× while keeping SLA violation rates below 9%, effectively validating the superiority of operator-level scheduling in multi-tenant scenarios.
📝 Abstract
Container-granularity scheduling leaves abundant short-lived idle slices within containers unexploited. Reallocating containers is too heavyweight to utilize such fine-grained opportunities under SLA constraints, and operator-level scheduling requires reasoning about dependencies, memory safety, and cluster-wide execution dynamics in real time. In this paper, we present SliceScheduler, a dynamic operator-level scheduling system for multi-tenant model serving. The key idea is to expose cluster-wide operator execution state and enable what-if reasoning over scheduling decisions. SliceScheduler consists of four key components. First, we introduce the Global Mapping Graph (GMG), a unified abstraction that captures operator dependencies, tensor shapes, resource mappings, and execution states, providing a real-time, cluster-wide view with explicit resource semantics. Second, we build a global simulator on top of GMG to predict operator-level execution and memory evolution under candidate placements. Third, we design an incremental, simulation-based scheduling module that selects placements to exploit fragmented idle slices while avoiding memory violations and preserving SLA. Finally, we develop an operator executor that materializes scheduling decisions on GPUs and coordinates computation and cross-accelerator transfers. We implement SliceScheduler as a PyTorch backend and evaluate it using production trace replay. Experimental results show that SliceScheduler improves token throughput by 1.10--2.29$\times$ compared to existing approaches, while maintaining SLA violations within 9\%. SliceScheduler demonstrates that operator-level scheduling is a practical and effective approach to improving GPU utilization for multi-tenant LLM serving.
Problem

Research questions and friction points this paper is trying to address.

Multi-tenant Model Serving
Operator-level Scheduling
GPU Utilization
Resource Fragmentation
SLA Constraints
Innovation

Methods, ideas, or system contributions that make the work stand out.

Operator-level Scheduling
Global Mapping Graph
Simulation-guided Scheduling
Multi-tenant Model Serving
What-if Reasoning