Reinforcement learning to choose optimizers

๐Ÿ“… 2026-09-01
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
็ ”็ฉถ้€š่ฟ‡ๅผบๅŒ–ๅญฆไน ้€‰ๆ‹ฉไผ˜ๅŒ–ๅ™จ็š„ๆ–นๆณ•๏ผŒ่งฃๅ†ณไธๅŒ้—ฎ้ข˜ๅ’Œ้˜ถๆฎตไธ‹ๆœ€ไผ˜ไผ˜ๅŒ–ๆ–นๆณ•็š„้€‰ๆ‹ฉ้—ฎ้ข˜๏ผŒ้‡‡็”จๅบๅˆ—ๅ†ณ็ญ–ๅˆถๅฎš็ญ–็•ฅไปฅ้€‚ๅบ”ๆ€ง้€‰ๆ‹ฉๆœ€ๅˆ้€‚็š„ไผ˜ๅŒ–ๅ™จใ€‚
๐Ÿ“ Abstract
No single optimization method is uniformly best for all problems, and the most suitable optimizer choice can change during a run. Existing approaches that change optimizer during execution typically predetermine part of the strategy: the portfolio is restricted to one algorithm class, the switch occurs once at a fixed time, or the frequency of decisions is treated as a hyperparameter rather than a learned one. We introduce "Reinforcement Learning to Choose Optimizers", which formulates the optimization algorithm choice as a sequential decision-making problem. At each decision, a recurrent policy reads the current run state and decides both which optimizer should be used next and for how long. The portfolio includes both gradient-based and derivative-free optimizers, and each switch passes on the current best solution and a representative step size. A context proxy conditions a gating network over expert heads, and training employs a decoupled actor-critic whose return is expressed in the same empirical runtime distribution metric used at evaluation. Training tasks and portfolio are designed jointly so that no optimizer dominates. On unseen problems, the learned policy outperforms every portfolio optimizer at all but the smallest budgets, and it remains robust under distribution shift.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Optimizers
Sequential Decision-Making
Gradient-based Optimizers
Derivative-free Optimizers
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Sequential Decision Making
Optimizer Portfolio
Gating Network
Decoupled Actor-Critic
๐Ÿ”Ž Similar Papers
2024-07-09Neural Information Processing SystemsCitations: 3
๐Ÿ’ผ Related Jobs
No related jobs found.
M
Martin van der Schelling
Department of Mechanical Engineering, Delft University of Technology, Delft, the Netherlands
D
Deepesh Toshniwal
Institute of Applied Mathematics, Delft University of Technology, Delft, the Netherlands
Miguel A. Bessa
Miguel A. Bessa
Associate Professor, Brown University
computational mechanicsmachine learningoptimization