Zero-shot Generalization in Inventory Management: Train, then Estimate and Decide
In real-world inventory management, dynamic uncertainty in demand and lead-time distributions severely limits the generalization capability of existing deep reinforcement learning (DRL) methods. To address this, we propose a novel “Train → Estimate → Decide” three-stage framework and introduce GC-LSN—the first zero-shot generalizable inventory agent for lost-sales settings with periodic demand and stochastic lead times. Our approach innovatively integrates a Super-MDP formulation with the Time-Evolving Distribution (TED) framework, combining nonparametric Kaplan–Meier distribution estimation, online parametric identification, and policy self-adaptation. Experiments demonstrate that GC-LSN significantly outperforms classical heuristics under known parameters and surpasses state-of-the-art online learning methods with worst-case guarantees under unknown demand and lead-time distributions. Crucially, it enables real-time, robust decision-making across unseen distributional shifts—achieving true zero-shot generalization in complex, non-stationary inventory environments.