Beyond Static Bias: Adaptive Multi-Fidelity Bandits with Improving Proxies

📅 2026-05-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of traditional multi-fidelity multi-armed bandit frameworks, which assume a fixed bias between low- and high-fidelity sources and thus fail to capture dynamic surrogates—such as large language models (LLMs)—that improve with use. We study an adaptive multi-fidelity setting where the low-fidelity source becomes increasingly informative over repeated usage, focusing on a two-fidelity scenario. To model the evolving approximation error, we introduce a selective average mismatch bound and devise a threshold-driven adaptive continuation strategy that reduces the high-fidelity evaluation cost for intermediate arms from logarithmic to bounded. Within an optimistic framework combining dynamic confidence bounds and threshold-based sampling, we establish instance-dependent regret upper bounds and demonstrate significant improvements in cost-weighted regret on both synthetic benchmarks and LLM-based policy evaluation tasks.
📝 Abstract
As an extension of the classical multi-armed bandit problem, multi-fidelity multi-armed bandits (MF-MAB) enable individual arms to be evaluated using diverse feedback sources that vary in both cost and accuracy. Prior stochastic models typically assume fixed low-to-high fidelity discrepancies, whereas modern proxy sources, such as learning-based simulators and Large Language Models (LLMs), can be improved using additional calibration. We investigate adaptive MF-MAB with improving proxy sources, and focus on the canonical two-fidelity case in which the low-fidelity source becomes more informative with repeated use. To capture this dynamic, we introduce a selected-average mismatch bound that converts dynamic low-fidelity observations into improvement-aware confidence bounds for the high-fidelity target. We propose the Threshold-Based Adaptive Continuation Companion (TACC), an optimistic algorithm that uses a bounded continuation rule to decide when low-fidelity sampling remains cost-effective and when to escalate. We prove an instance-dependent regret bound showing that, for detected intermediate arms, adaptive continuation replaces logarithmic high-fidelity confirmation with bounded low-fidelity continuation. Experiments on synthetic bandits and an LLM-as-a-judge policy-evaluation task examine when continuation improves cost-weighted regret.
Problem

Research questions and friction points this paper is trying to address.

multi-fidelity bandits
adaptive learning
improving proxies
dynamic bias
cost-effective evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

adaptive multi-fidelity bandits
improving proxies
TACC algorithm
selected-average mismatch bound
cost-weighted regret
🔎 Similar Papers
2024-06-05Neural Information Processing SystemsCitations: 1
💼 Related Jobs
No related jobs found.
M
Muyun Lu
Department of Industrial and Systems Engineering, University of Houston
H
Haoyang Hong
School of Electrical Engineering and Computer Science, Oregon State University
Huazheng Wang
Huazheng Wang
Assistant Professor, Oregon State University
Reinforcement LearningMachine LearningInformation Retrieval
Ying Lin
Ying Lin
University of Houston
statistical learningbiomedical data analysismedical decision making