科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ International Journal of Progressive Research in Engineering Management and Science2026-07-31· Tardiness

DIGITAL-TWIN-EMBEDDED DEEP REINFORCEMENT LEARNING FOR ADAPTIVE SCHEDULING AND REPLANNING IN SMART MANUFACTURING UNDER COMPOUND DISTURBANCE

原始摘要(英文原文)· Original abstract
Smart manufacturing environments are increasingly dynamic and uncertain, with stochastic job arrivals, unplanned machine breakdowns and shifting order priorities placing continuous pressure on production scheduling.Classical and metaheuristic scheduling methods offer no continuous adaptation once a disturbance occurs, existing digital-twin scheduling systems improve disturbance perception but leave the rescheduling decision itself threshold-or ruletriggered, and current deep-reinforcement-learning (DRL) scheduling studies validate hand-crafted, single-disturbance policies on one shop-floor configuration only.This paper proposes an adaptive scheduling and replanning framework that embeds a Proximal Policy Optimization (PPO) agent, conditioned on a graph-neural-network (GNN) state encoder and a digital-twin-supplied belief state, directly inside the decision-and-control layer of a six-part smartmanufacturing architecture.The scheduling problem is formalised as a Markov Decision Process / Partially Observable Markov Decision Process (MDP/POMDP) that jointly represents job-arrival, machine-breakdown and priority-change disturbances, individually and in combination, and the agent is trained across varying shop-floor configurations and disturbance combinations rather than a single fixed setting.The proposed agent was benchmarked in a discrete-event shop-floor simulator against a classical priority-dispatching baseline, a digital-twin rule-triggered pipeline, and four representative DRL scheduling models re-implemented under a shared protocol.Across single-and compound-disturbance scenarios the proposed agent achieved the lowest makespan and tardiness and the highest machine utilisation of all seven models, degraded the least under compound disturbance, generalised to unseen shopfloor scales with the smallest performance gap, remained within a real-time-adequate decision latency, and produced trace-level, auditable justifications for the large majority of sampled dispatch decisions, a capability absent from every reviewed comparator.An ablation study attributes these gains jointly to the graph encoder, the belief-state mechanism, and the multi-disturbance training regime rather than to the learning algorithm alone.All reported results are simulation-based; sim-to-real transfer and operator-in-the-loop evaluation are identified as future work.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

DIGITAL-TWIN-EMBEDDED DEEP REINFORCEMENT LEARNING FOR ADAPTIVE SCHEDULING AND REPLANNING IN SMART MANUFACTURING UNDER COMPOUND DISTURBANCE — 科研速览 Science Skim