Opportunity Is Not Realizability: Selection-Valid Diagnostics for Multi-LLM Routing

📅 2026-08-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the overestimation of deployable router performance in existing large language model (LLM) routing evaluation methods, which suffer from selection bias and information leakage. The authors propose the first selection-valid diagnostic framework that explicitly disentangles three distinct objectives: opportunity, attainability, and realized gain, and constructs post-selection valid confidence intervals. Their approach innovatively integrates a signal information sandwich theorem, Bayesian optimal gain analysis, greedy pool construction under submodular coverage, and multiple hypothesis testing correction. Experiments across four benchmarks reveal that the actual attainable gain constitutes only 7.5%–14.4% of the theoretical opportunity, and while the strongest routers outperform any fixed model, the majority of potential gains remain unrealizable in practice.
📝 Abstract
Oracle routing measures how much a pool of language models could gain from per-query selection, but the diagnostic has two flaws: testing against a best fixed model selected on the same examples invalidates paired inference, and a full-information oracle sees outcomes no deployable router observes. We separate three estimands (outcome-oracle opportunity, the Bayes-optimal gain from a declared pre-answer signal, and the held-out gain of a learned router) and prove selection-valid confidence intervals that survive choosing the best fixed model or the best member of a router family, a signal-information sandwich, and a $(1-1/e)$ greedy guarantee for building compact pools from submodular complementary coverage. On eight checkpoints from six families over four benchmarks, selection-valid intervals certify a population oracle gap of $9.7$--$30.7$ points on every task, yet the strongest deployable prompt router recovers only $7.5$--$14.4\%$ of it, and the simultaneous interval for the best of eleven tested policies has lower limit zero throughout. The realizable share of oracle opportunity is small and certifiable: strong routers beat the best fixed model, and most of the gap remains.
Problem

Research questions and friction points this paper is trying to address.

oracle routing
selection validity
multi-LLM routing
realizability
diagnostic evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

selection-valid inference
multi-LLM routing
oracle opportunity
submodular coverage
Bayes-optimal routing
💼 Related Jobs
No related jobs found.
Ibne Farabi Shihab
Ibne Farabi Shihab
Iowa State University
Deep LearningroboticsLarge Language Model
A
Abu Sa-Adat Mohamed Moon-Im Al Ahsan
Department of Computer Science & Engineering, BRAC University
M
Md Najmus Swaqeeb
Department of Computer Science & Engineering, BRAC University