Muhamed Ramees Cheriya Mukkolakkal
Current microservice auto-scaling solutions operate in isolation, focusing on individual service metrics without considering global cloud resource availability, cross-datacenter performance, or mission-critical application priorities. This paper presents InfraLLM, a novel framework leveraging large language models to orchestrate intelligent, context-aware auto-scaling decisions across entire cloud infrastructures. Our approach integrates three key components: a distributed Collection Service for comprehensive metric aggregation, an LLM Service for predictive resource allocation, and an Execution Service for policy enforcement. Evaluation across large-scale Kubernetes deployments demonstrates up to 57.2% reduction in CPU over-utilization, 51.1% improvement in resource allocation efficiency, 48% reduction in average response time, and 16× reduction in SLO violations compared to traditional per-service auto-scaling approaches. InfraLLM represents a paradigm shift from reactive, service-level scaling to proactive, infrastructure-wide resource orchestration.