Meta-LinEXP3: Online-within-Online Learning for Adversarial Linear Contextual Bandits

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为了解决随机动作集对抗线性情境盗匪问题,提出Meta-LinEXP3算法,通过先前任务构建任务级先验以指导内层学习者,并在已知与未知分布情况下分别设计估计器来优化遗憾界。
📝 Abstract
Meta-learning has emerged as an effective paradigm for transferring knowledge across sequential bandit tasks. While substantial progress has been made for stochastic bandits and non-contextual adversarial bandits, meta-learning for adversarial linear contextual bandits (ALCBs) with random action sets remains largely unexplored. To address this problem, we propose Meta-LinEXP3, an online-within-online algorithm that constructs a predictable task-level prior from completed tasks to guide the inner LinEXP3 learner. For known context distributions, we develop a policy-centered estimator that achieves an intrinsic-dimension $\mathcal{O}(\sqrt{n})$ per-task regret bound. For unknown distributions, we introduce a past-only regularized moment estimator with an $\mathcal{O}(n^{2/3})$ leading regret term and explicit finite-sample error. We further establish a direct connection between prior accuracy and transfer regret, showing that increasingly accurate priors yield sublinear transfer-dependent regret across tasks. Experiments demonstrate the effectiveness of Meta-LinEXP3, including its application to structured hyperspectral tensor sampling.
Problem

Research questions and friction points this paper is trying to address.

Meta-learning
Adversarial Linear Contextual Bandits
Random Action Sets
Innovation

Methods, ideas, or system contributions that make the work stand out.

Meta-LinEXP3
adversarial linear contextual bandits
online-within-online learning
policy-centered estimator
transfer regret
🔎 Similar Papers
No similar papers found.
H
Hao Li
College of Science, National University of Defense Technology, Changsha 410073, China
Jie Xu
Jie Xu
Argonne National Lab Stanford university
Z
Zheng Xie
College of Science, National University of Defense Technology, Changsha 410073, China