Shared KV Caching for Replicated 27B Inference: Correctness Failures and Performance Boundaries

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
📝 Abstract
Shared host-memory caching can avoid repeated prefill when a request moves between inference replicas. Its usefulness depends on both correct state transfer and lost prefix locality. We study two single-GPU 27B vLLM replicas sharing a 256 GiB LMCache pool. After adopting an existing packed-page patch, we isolate a raw-pointer fallback that omits the dependency on the current CUDA stream. Controlled byte tests fail under an imposed delay and pass when the dependency is restored; the existing mixed allocator provides a working deployment path. Full-pool allocation checks and service regression complete the validation. A four-block OFF-ON-ON-OFF comparison contains 768 measured requests within two block pairs. Median cross-replica time to first content token falls from 31.715 to 0.605 seconds at 128k input and from 92.047 to 0.790 seconds at 256k. Six-turn synthetic sessions alternating replicas improve by approximately 35% and 45% at initial contexts of 32k and 128k, while fixed placement shows little benefit. This engineering case study identifies practical validation steps and the locality conditions in which shared caching pays off.
Problem

Research questions and friction points this paper is trying to address.

Shared KV Caching
Correctness Failures
Performance Boundaries
Inference Replicas
Innovation

Methods, ideas, or system contributions that make the work stand out.

Shared KV Caching
Replicated Inference
LMCache
CUDA Stream Dependency
Performance Optimization
🔎 Similar Papers
No similar papers found.