科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ International Journal of Scientific Research in Computer Science Engineering and Information Technology2026-08-01· Cloud computing

SAFE-HealCloud: Safety-Aware, Agentic Self-Healing for Cloud Infrastructure

Prudvi Saisaran Ponduru, Pavani Priya Vyshnavi Nandanavanam, Sai Kesav Kumar Ponduru

原始摘要(英文原文)· Original abstract
Cloud infrastructure failures are increasingly difficult to detect, diagnose, and remediate because production environments combine microservices, Kubernetes control loops, service meshes, serverless workloads, infrastructure-as-code, continuous delivery, and heterogeneous telemetry. Reactive monitoring and manual incident response remain necessary, but they do not scale to the volume, velocity, and causal complexity of modern cloud operations. This paper provides a structured synthesis of scholarly, industry, and standards-based work on AI-driven self-healing for cloud infrastructure, with emphasis on AIOps, AgentOps, LLM-based cloud operations, anomaly detection, causal root cause analysis, graph learning, reinforcement learning, automated remediation, Kubernetes self-healing, observability, chaos engineering, and self-healing infrastructure-as-code. We propose SAFE-HealCloud, a safety-aware, agentic, feedback-driven framework that integrates telemetry ingestion, multimodal observability, anomaly detection, failure prediction, causal RCA, retrieval-augmented LLM reasoning, policy guardrails, risk-scored remediation planning, controlled execution, verification, rollback, human approval, and continuous learning. A formal model defines cloud state, observability vectors, failure states, action spaces, remediation policies, rewards, constraints, and reliability objectives. We also specify reproducible Kubernetes-based evaluation designs, metrics, algorithms, risk controls, and operational use cases. The central finding is that near-term practical value lies in graduated autonomy: low-risk reversible actions can be automated, medium-risk actions should be canaried and policy-gated, and high-risk changes should remain human-approved. Designed in this way, AI-driven self-healing can reduce detection and recovery times, preserve error budgets, improve operator productivity, and strengthen digital resilience without sacrificing safety, auditability, or governance.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

SAFE-HealCloud: Safety-Aware, Agentic Self-Healing for Cloud Infrastructure — 科研速览 Science Skim