Cong Zhang, Fan Wu, Zhenyu Wang, Xiaohan Wang, Huadóng Ma, Yuanan Liu
As cloud computing revolutionizes various fields, the demand for scalable and flexible computing resources has grown significantly. Applications in large-scale engineering simulations, artificial intelligence, and data analysis require substantial computational power and memory, placing pressure on traditional systems. With its heterogeneous resource pools, cloud computing offers a promising solution by enabling the dynamic allocation of diverse, distributed resources. These resource pools facilitate parallel task execution, accelerate computations, and enhance system flexibility. However, efficiently managing these resources remains a complex, NP-hard challenge due to the vast search space, resource fragmentation, and the need for dynamic adjustments. In this paper, we first develop a novel multi-task flow representation model using a deep graph neural network (GNN) and a resource pool representation model based on a convolutional neural network (CNN). These models describe the dependencies among multiple tasks, resource requirements, and resource distribution in the multi-DAG job arrival scenario. Then, we design a dynamic resource allocation strategy model based on deep reinforcement learning (DRL) to reduce job processing time. Finally, we compare the performance of our proposed method with the other eight baseline algorithms using test datasets of various scales and under different arrival modes consisting of Montage, CyberShake, Broadband, Epigenomics, LIGO, VGG 16-SVD, and Edge Detection datasets.