Bandit Submodular Maximization under Matroid Constraints: Learning Compressed Exchange Policy

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究了在拟阵约束下对抗性带状子模最大化问题,通过学习压缩交换策略,提出了一种具有次线性遗憾界的多项式时间算法。
📝 Abstract
We study adversarial bandit maximization of monotone submodular functions under a matroid constraint. For a rank-$k$ matroid on $n$ elements, we give a randomized oracle-polynomial algorithm that makes one feasible value query per round and has expected $(1-1/e)$-regret $\widetilde O(n^{1/3}k^{2/3}T^{2/3})$. This is the first sublinear-regret algorithm for adversarial bandit submodular maximization under general matroid constraints. Technically, we view the problem as learning an exchange policy for the Poisson base walk. This connects the problem to contextual bandits and gives an information-theoretic sublinear-regret guarantee, but directly learning the exponentially many policies requires exponential time and space. We therefore introduce \emph{balanced fractional exchanges}, which compress the policy mixture into a single fractional base while retaining the exchange information needed by the Poisson analysis. This leads to an polynomial time algorithm with the same regret guarantee.
Problem

Research questions and friction points this paper is trying to address.

adversarial bandit
submodular maximization
matroid constraint
Innovation

Methods, ideas, or system contributions that make the work stand out.

adversarial bandit submodular maximization
matroid constraints
balanced fractional exchanges
contextual bandits
🔎 Similar Papers
2024-10-02International Conference on Machine LearningCitations: 1