Risk-Aware and Stable Edge Server Selection Under Network Latency SLOs
This work addresses the challenge of dynamic server selection in edge computing under stringent latency SLOs, where decisions must jointly manage tail-risk violations and switching stability. The authors propose a lightweight, interpretable decision framework that uniquely co-optimizes tail-risk control and switching stability: it estimates SLO violation risk using normal approximation and the Cantelli inequality, while incorporating a hysteresis mechanism to suppress excessive switching. Experimental results under a 0.5-second SLO demonstrate that, compared to a baseline relying solely on mean latency, the proposed method reduces deadline miss rate from 39% to 34%, cuts switching frequency by 88% (down to 5.5%), and maintains average latency stably at 0.45 seconds, thereby significantly enhancing both system robustness and efficiency.