Kai Peng, Xudong Liu, Di Han, Yi Hu, Menglan Hu, Chao Cai, Zehui Xiong
The rapid development of AI accelerates the implementation and delivery of AI applications in diverse fields. In cloud-edge collaboration, delivering a complete AI application relies on the robust coordination between AI-supporting microservices and AI services. However, most existing studies only coarsely considered monolithic AI service orchestration while neglecting microservice orchestration. Such coarse-grained orchestration severely impacts application performance. To enable diverse high-performance AI applications, fine-grained hybrid orchestration of AI services and microservices (HOAIM) is highly desirable, yet presents formidable challenges. Due to heterogeneous services, call dependencies, and service multiplexing, fine-grained hybrid orchestration modeling is highly non-trivial. Moreover, the tight coupling between deployment and routing results in a complex joint optimization problem. To address this, we first propose a heterogeneous service orchestration network that supports orchestration optimization and automated management. Then, based on queuing networks and multi-instance models, we conduct an accurate analysis of delay and load. Furthermore, to achieve efficient hybrid orchestration, we propose preference-driven resource allocation and instance computation algorithms, along with reinforcement learning with action masking and reward shaping. Finally, extensive trace-driven simulations demonstrate that our algorithms optimize average response delay by up to 41.83%, and achieve significant advantages in load balancing, response success rate, and resource efficiency.