OR-Transformer: Scaling Real-Time Decision-Making to 1,000 Items

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入OR-Transformer,一种基于深度强化学习和变换器架构的方法,解决了大规模供应链中实时决策问题,特别是在处理高达1024个库存项目时,显著优于现有方法。
📝 Abstract
Modern supply chain operations can require coordinating replenishment across thousands of heterogeneous items under correlated stochastic demand, heterogeneous lead times, and shared fixed ordering costs, yielding observation spaces exceeding $10^4$ dimensions. At this scale, rolling-horizon stochastic mixed-integer linear programs (MILPs) become prohibitively slow, while standard reinforcement learning (RL) methods face increasingly challenging credit assignment in high-dimensional action spaces. We introduce OR-Transformer, a deep reinforcement learning framework for joint replenishment under stochastic demand, with an item-permutation-equivariant Transformer architecture and pathwise-gradient training through the inventory dynamics. Across problem sizes up to 1,024 inventory items, OR-Transformer increasingly outperforms learning-based and rolling-horizon MILP baselines as scale grows. It also reduces online decision-making time by over 4 million times relative to MILP solvers, enabling real-time, large-scale deep RL in supply chain operations.
Problem

Research questions and friction points this paper is trying to address.

real-time decision-making
supply chain operations
stochastic demand
high-dimensional action spaces
reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

OR-Transformer
item-permutation-equivariant Transformer
pathwise-gradient training
real-time decision-making
large-scale supply chain operations
S
Shuze Daniel Liu
Massachusetts Institute of Technology
D
David Simchi-Levi
Massachusetts Institute of Technology
Claire Chen
Claire Chen
PhD student, Stanford University
contact-rich manipulationrobot learningmulti-modal sensing
C
Chutong Gao
Massachusetts Institute of Technology
Shangtong Zhang
Shangtong Zhang
University of Virginia
reinforcement learningstochastic approximation