Systematically diagnose and fix a broken Kubernetes cluster with multiple interconnected failures
Diagnose OOMKilled pods and fix resource configuration by analyzing actual usage patterns
Configure horizontal pod autoscaling, right-size resource requests, and add resilience for services under load