Error-Aware Reverse Auction Mechanism for Large Language Model Routing
This work addresses the limitations of centralized prediction-based routing in large language models (LLMs), which often leads to misaligned information risks and scalability bottlenecks. To overcome these issues, the paper introduces, for the first time, a reverse auction mechanism into LLM routing and proposes the error-aware EA-RAM framework. In this framework, model providers autonomously bid their success rates and costs, while explicit modeling of dual sources of noise—arising from both prediction and evaluation—enables robust handling of uncertainty. The mechanism is shown to be Bayesian incentive-compatible and individually rational, with a provable upper bound on social welfare loss. Empirical results demonstrate that EA-RAM consistently outperforms centralized baselines across both simulated and real-world benchmarks, maintaining robustness under dual-error conditions and significantly advancing the cost-performance Pareto frontier.