🤖 AI Summary
Combinatorial optimization problems pose significant challenges for deep reinforcement learning (DRL) due to their discrete, exponentially large solution spaces, often leading DRL to premature convergence to local optima. While genetic algorithms (GAs) offer strong global exploration, they suffer from low sample efficiency. To bridge this gap, we propose the Evolutionary Augmentation Mechanism (EAM)—a plug-and-play framework enabling dynamic, closed-loop coupling of DRL and GA during training. EAM integrates policy-based sampling, customized genetic operators, and solution re-injection to jointly optimize policy and solution distributions. Crucially, we derive an upper bound on the KL divergence between the policy distribution and the evolutionary solution distribution, theoretically guaranteeing distributional stability. Compatible with mainstream DRL architectures (e.g., Attention Model, POMO) and GA operators, EAM achieves state-of-the-art performance on TSP, CVRP, PCTSP, and OP benchmarks—simultaneously improving convergence speed and solution quality.
📝 Abstract
Combinatorial optimization problems are notoriously challenging due to their discrete structure and exponentially large solution space. Recent advances in deep reinforcement learning (DRL) have enabled the learning heuristics directly from data. However, DRL methods often suffer from limited exploration and susceptibility to local optima. On the other hand, evolutionary algorithms such as Genetic Algorithms (GAs) exhibit strong global exploration capabilities but are typically sample inefficient and computationally intensive. In this work, we propose the Evolutionary Augmentation Mechanism (EAM), a general and plug-and-play framework that synergizes the learning efficiency of DRL with the global search power of GAs. EAM operates by generating solutions from a learned policy and refining them through domain-specific genetic operations such as crossover and mutation. These evolved solutions are then selectively reinjected into the policy training loop, thereby enhancing exploration and accelerating convergence. We further provide a theoretical analysis that establishes an upper bound on the KL divergence between the evolved solution distribution and the policy distribution, ensuring stable and effective policy updates. EAM is model-agnostic and can be seamlessly integrated with state-of-the-art DRL solvers such as the Attention Model, POMO, and SymNCO. Extensive results on benchmark problems (e.g., TSP, CVRP, PCTSP, and OP) demonstrate that EAM significantly improves both solution quality and training efficiency over competitive baselines.