Semantic Bandits: In-Context Exploration-Exploitation is Biased by Semantic Priors
This study addresses the exploration-exploitation imbalance caused by semantic priors in large language model decision-making by proposing a Semantic Multi-Armed Bandit framework. This approach formally quantifies how the alignment between linguistic labels and reward structures influences exploration strategies. Through in-context learning and inductive bias analysis, we reveal the bias effects inherent in label semantics and reward signals. Empirical results demonstrate that semantically consistent labels significantly enhance decision performance, whereas mismatches cause severe degradation. Furthermore, negative rewards elicit greater exploration than positive ones, validating scale biases present in pretraining data. Collectively, this work offers novel insights into understanding and optimizing the decision-making behaviors of large language models.