科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ International Journal of Industrial Engineering Computations2026-01-01· Reinforcement learning

A dynamic inventory optimization model for supply chains by deep reinforcement learning

Shengyi Zhou, Lili Zhou, Liang Chen

原始摘要(英文原文)· Original abstract
This study proposes a dynamic inventory optimization model based on deep reinforcement learning (DRL) to improve the intelligence of inventory management in multi-node supply chain systems, aiming to mitigate stockout risks and reduce inventory holding costs. The model is built upon three key components. First, a simulation environment is developed to mirror real-world supply chain operations, incorporating multiple products, warehouses, and sales channels. The simulator takes into account critical factors such as lead times, demand variability, and order fulfillment constraints. Second, the Soft Actor-Critic (SAC) algorithm is employed as the core of the learning strategy. To enhance the model’s adaptability to dynamic inventory changes, a dual-stage state representation and an attention-enhanced perception mechanism are introduced. Third, the model is trained and validated using the publicly available Instacart Market Basket Analysis dataset, which serves as a benchmark platform for testing retail replenishment strategies. To assess its effectiveness, the proposed SAC-based model is compared with several baseline methods, including Deep Deterministic Policy Gradient (DDPG), Proximal Policy Optimization (PPO), and an LSTM-based Model Predictive Control approach (LSTM-MPC). Experimental results across six representative scenarios show that the SAC model consistently outperforms competing approaches. Specifically, in high-demand volatility and limited-capacity scenarios, the SAC model reduces average inventory costs to ¥69,900 and ¥97,200, respectively, substantially lower than those achieved by DDPG (¥120,700 and ¥136,300). Under compound disturbance conditions, SAC limits the number of stockout events to 59, which is 50% fewer than the rule-based benchmark, significantly improving service level performance. Moreover, in high-frequency, short-cycle environments, SAC achieves the lowest inventory variability, with a standard deviation of just 4.12, indicating superior policy stability. Overall, the proposed DRL-based model exhibits strong robustness and adaptability, offering a practical solution for intelligent and resilient inventory control. This study highlights the potential of advanced reinforcement learning methods in addressing complex scheduling problems in real-world supply chain management.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A dynamic inventory optimization model for supply chains by deep reinforcement learning — 科研速览 Science Skim