Quality-Constrained Routing over a Fixed Pool of Quantized Mixture-of-Experts Instances

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究在固定池内通过质量约束路由解决量化混合专家模型实例的请求分配问题,提出Fragility-Weighted Perplexity方法以优化吞吐量和质量。
📝 Abstract
Quantized Mixture-of-Experts (MoE) services can hold several pre-materialized instances of one base model, but quantization damage varies sharply across requests and bitwidths. Because instance materialization and replica counts consume memory and require slow reconfiguration, we treat them as upstream provisioning decisions and study routing within a fixed resident pool. Within this fixed-pool boundary, we route each request to maximize modeled throughput under a class-level expected quality-degradation budget and measured instance capacities. To predict this request-specific risk, we introduce FWP (Fragility-Weighted Perplexity), computed from prompt tokens on a reference-instance prefill and calibrated to candidate-instance degradation. Underlying FWP is an exact two-expert affinity--fragility decomposition and a conditional multi-layer top-$k$ expansion whose bias, interaction, route-change, separability, and higher-order terms remain explicit. Using these calibrated risks, a window-level linear program yields a signed reduced-reward score that is KKT-consistent with the LP optimum under optimal prices and primal-feasible tie allocation. On 88 extended Qwen prompts, complete W2, W3, and W4 instances quantizing all 6,144 expert blocks incur mean $Δ$NLL of $0.9437$, $0.1832$, and $0.0513$. Under the same population and $τ=0.1513$, FWP allocation reaches a $1.284\times$ offline model-based multiplier versus $1.253\times$ for request-agnostic mixing and $1.000\times$ for static W4, an incremental $2.5\%$ relative FWP gain.
Problem

Research questions and friction points this paper is trying to address.

Quantized Mixture-of-Experts
Quality-Degradation Budget
Routing
Throughput Maximization
Instance Materialization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Quantized Mixture-of-Experts
Fragility-Weighted Perplexity
Linear Programming
Routing Strategy
Quality-Degradation Budget
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhenghong Huang
The Hong Kong University of Science and Technology
H
Hongfan Wu
The Hong Kong University of Science and Technology
Jiheng Zhang
Jiheng Zhang
The Hong Kong University of Science and Technology
Applied ProbabilityStochastic Modeling and OptimizationNumerical Methods and Algorithm