一键导入
kubernetes-debugging
Kubernetes debugging for pod failures and networking.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Kubernetes debugging for pod failures and networking.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Transforms a solidified idea, spec file, design doc, exploration output, or grilling session output (from /idea, /grill-me, or /grill-with-docs) into a structured, ticket-ready implementation plan. Use this skill whenever the user wants to turn any requirement, feature idea, spec, or decision into a detailed plan — even if they just say "plan this", "make a plan for", "let's plan out", or pastes raw notes and asks what to build. Always trigger this skill when the input looks like a feature request, system design, product requirement, or grilling output that needs to be broken down into actionable steps. If the user wants to go from idea → plan → build, this is the skill for the plan step.
Create new agent skills with proper structure, progressive disclosure, and bundled resources. Use when user wants to create, write, or build a new skill.
Audit, plan, and implement SEO improvements across technical SEO, on-page optimization, structured data, Core Web Vitals, and content strategy. Use when the user wants better search visibility, SEO remediation, schema markup, sitemap/robots work, or keyword mapping.
notifies user about status of today (TODOs, emails, repos tracking, news, etc.)
Turn the current conversation context into a PRD. Use when user wants to create a PRD from the current context.
| name | kubernetes-debugging |
| description | Kubernetes debugging for pod failures and networking. |
| routing | {"triggers":["kubernetes debug","pod failure","pod crashloop","kubectl logs","OOMKilled","pod pending"],"category":"kubernetes","pairs_with":["kubernetes-security","service-health-check"]} |
Systematic diagnosis of pod failures, networking issues, and resource problems using a structured triage flow: describe, logs, events, exec.
| Signal | Reference | Size |
|---|---|---|
| CrashLoopBackOff, OOMKilled, config error, health check, liveness probe, ImagePullBackOff, image pull, registry auth, Pending, FailedScheduling, node affinity, taint, PVC | references/crash-diagnosis.md | ~140 lines |
| service resolution, DNS, nslookup, CoreDNS, port-forward, NetworkPolicy, ingress, egress | references/network-debugging.md | ~50 lines |
| CPU throttling, memory limit, OOMKill, ephemeral storage, DiskPressure, debug container, distroless, kubectl reference, rollout, exec | references/resource-debugging.md | ~100 lines |
Load greedily. If the user's question touches any signal keyword, load the matching reference before responding. Multiple signals matching = load all matching references.
Follow this sequence for every pod or workload issue. Do not skip steps -- many failures (scheduling, image pull, volume mount) are only visible in events and describe output, not in logs, so jumping straight to logs misses them.
Always specify -n <namespace> explicitly in every command; never rely on the default context namespace, because the wrong namespace silently returns empty or misleading results.
# 1. Get an overview of the resource state
kubectl get pods -n <namespace> -o wide
# 2. Describe the resource for events, conditions, and status
kubectl describe pod <pod-name> -n <namespace>
# 3. Check current container logs
kubectl logs <pod-name> -n <namespace> -c <container-name>
# 4. Check previous container logs (critical for CrashLoopBackOff)
# Always check --previous before current logs for crashed containers,
# because deleting or restarting the pod destroys these logs permanently.
kubectl logs <pod-name> -n <namespace> -c <container-name> --previous
# 5. Check namespace events sorted by time
kubectl get events -n <namespace> --sort-by='.lastTimestamp'
# 6. If the container is running, exec in for live inspection
kubectl exec -it <pod-name> -n <namespace> -c <container-name> -- /bin/sh
Use read-only commands (describe, logs, get) to gather evidence before proposing any modifications. Never suggest changes based on assumptions -- gather diagnostic output first.
Based on triage output, load the appropriate reference and follow its diagnosis flow:
| Symptom | Reference |
|---|---|
| Pod status CrashLoopBackOff, ImagePullBackOff, or Pending | references/crash-diagnosis.md |
| Service unreachable, DNS failure, connection refused | references/network-debugging.md |
| CPU throttling, OOMKill, disk pressure, need debug container | references/resource-debugging.md |
Cause: The Service selector does not match any running pod labels.
Solution: Compare kubectl get svc <name> -o yaml selector with kubectl get pods --show-labels. Fix the label mismatch.