科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Computers & Operations Research2026-02-11· Computer science

Enhancing multi-agent deep reinforcement learning for flexible job-shop scheduling through constraint programming

Alexandre Jesus, Arthur Jorge Pereira Corrêa, Miguel Vieira, Catarina M. Marques, Cristóvão Silva, Samuel L Moniz

原始摘要(英文原文)· Original abstract
This paper introduces PRISMA , a hybrid multi-agent Deep Reinforcement Learning (DRL) framework for solving the Flexible Job-shop Scheduling Problem (FJSP). It uses Constraint Programming (CP) solutions to pretrain decentralized policies and to guide exploration during training. Although DRL can generate fast solutions for large combinatorial problems, it often fails to match the quality of optimization methods, motivating the integration with hybrid frameworks. The growing interest in embedding domain knowledge into learning algorithms has produced several hybrid formulations, yet their potential remains underexplored, particularly in multi-agent settings. PRISMA combines supervised and reinforcement learning within a multi-agent framework, where CP solutions are used to (i) learn expert decisions through imitation learning, and (ii) train an auxiliary network that guides DRL training via reward shaping. A shared graph network is adopted for transferring system-level knowledge into machine-level observations, enabling fast and consistent inference from enriched local embeddings. To the best of our knowledge, PRISMA introduces the first expert-derived guidance mechanism for the FJSP and is among the earliest to apply imitation learning within a multi-agent formulation. By combining both modules, it strengthens the bridge between optimization and learning-based methods, where such dual integrations remain scarce. Experimental results show faster convergence and higher solution quality than state-of-the-art DRL models. PRISMA achieves an average optimality gap of 6.74%, corresponding to a 50% relative improvement over the single-agent baseline, while reducing inference time. These findings reinforce the value of merging optimization accuracy with the flexibility of multi-agent DRL for efficient scheduling.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Enhancing multi-agent deep reinforcement learning for flexible job-shop scheduling through constraint programming — 科研速览 Science Skim