Skip to main content

kpanda-cluster-diagnosis

Use when a user asks to diagnose, inspect, or troubleshoot the health of a Kubernetes cluster managed by the DCE kpanda module. Symptoms include cluster unavailability, node NotReady, pending or failed Pods, and requests to check cluster status.

Jump to install

Source facts

Repository
DaoCloud/ai-skills-bench
Last source activity
September 7, 2026 at 09:29
Detected SKILL.md language
English
Stars
0
Forks
1

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.

Showing SKILL.md

SKILL.md
Source instructions ยท Read-only preview
name
kpanda:cluster-diagnosis
description
Use when a user asks to diagnose, inspect, or troubleshoot the health of a Kubernetes cluster managed by the DCE kpanda module. Symptoms include cluster unavailability, node NotReady, pending or failed Pods, and requests to check cluster status.
# Kpanda Cluster Diagnosis Diagnose cluster health through a standardized 4-step inspection workflow. **REQUIRED SUB-SKILL:** Use `dce` for all command execution, auth checks, and catalog discovery. ## Workflow ### Step 1 โ€” Cluster Overview - `dce container-management cluster get-cluster --name <cluster> -o json` - Verify cluster exists and status is Running. If not, report immediately. ### Step 2 โ€” Node Health - `dce container-management core list-nodes --cluster <cluster> -o json` - Flag NotReady, Cordoned, or pressured nodes. Continue regardless. ### Step 3 โ€” Abnormal Pod Discovery - `dce container-management core list-pods --cluster <cluster> -o json` - Find Pods not in Running/Succeeded. Collect by namespace. If none, skip Step 4. ### Step 4 โ€” Deep Diagnosis - `dce container-management core list-cluster-events --cluster <cluster> -o json` - `dce container-management core get-pod --cluster <cluster> --namespace <ns> --name <pod> -o json` - Correlate events with Pod states to infer root cause. ## User omitted cluster name Run `dce container-management cluster list-clusters -o json`, present list, ask user to pick one. ## Auth not established Stop and instruct user to run `dce auth login --hostname <host>`. ## Output Format Present a concise report in this order: 1. **Cluster State** โ€” name, status, version, provider 2. **Node Summary** โ€” total, Ready count, anomalies 3. **Pod Anomalies** โ€” count by non-Running phase, top affected namespaces 4. **Root-Cause Hypothesis** โ€” inferred from events + Pod states 5. **Recommended Action** โ€” one or two concrete next steps ## Rules - Prefer `-o json` for machine-readable output. - Do not guess flags or body shape. Confirm with `dce commands show` before executing unfamiliar commands. - Report empty API responses as "no resources found" rather than silently skipping. - Do not perform remediation (restart, delete, scale). This skill is read-only.
View on GitHub