科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ IEEE Transactions on Smart Grid2025-11-10· Computer science

Constrained Semi-MDP Formulation and Perception-Enhanced Safe Policy Learning for Efficient Dynamic Task Scheduling of Data Centers

Yiling Zhang, Yujian Ye, Jianxiong Hu, Heng Hu, Xi Zhang, Dezhi Xu

原始摘要(英文原文)· Original abstract
With the exponential growth of global data generation and the energy consumption of data centers (DCs), efficient dynamic task scheduling has become critical to optimize energy costs while ensuring the quality of the service. Model predictive control (MPC) method is burdened with significant computational complexity and exhibits limited uncertainty adaptability, heuristic methods overlook the energy cost optimization objective and may lead to sub-optimal schedules, while conventional reinforcement learning (RL) methods based on Markov decision processes (MDPs) face significant dimensionality challenge in discrete action spaces and struggle to satisfy the temporal-coupling task deadline constraints. This paper proposes a Constrained Semi-MDP (CSMDP) formulation for dynamic task scheduling considering heterogeneous tasks and servers. It introduces temporally-extended sub-actions executable across multiple sub-processes, omitting exhaustive search for appropriate task-server pair combinations. A dual-agent architecture involving a task selection agent and a server assignment agent are designed to work collaboratively to optimize energy costs while satisfying task deadlines and resource constraints. A safe RL method is proposed that employs a constraint-value network and a shielding mechanism for task selection and server assignment agents to enforce resource capacity and time-coupling deadline constraints, while jointly optimizing the latter with an action-value network to improve energy cost efficiency. Case studies based on Alibaba real production dataset validate the effectiveness of the proposed method in terms of scheduling policy optimality and safety, uncertainty generalization and computational efficiency, against state-of-the-art methods.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Constrained Semi-MDP Formulation and Perception-Enhanced Safe Policy Learning for Efficient Dynamic Task Scheduling of Data Centers — 科研速览 Science Skim