| name | rca |
| description | Root cause analysis workflows - systematic investigation of failures |
flowchart TD
FAIL([Failure]) --> RCA{"/rca"}
RCA -->|CI failure, no cluster| RCACI["rca:ci"]:::rca
RCA -->|HyperShift available| RCAHS["rca:hypershift"]:::rca
RCA -->|Kind available| RCAKIND["rca:kind"]:::rca
RCACI -->|Inconclusive| NEED{"Need cluster?"}
NEED -->|Yes| RCAHS
NEED -->|Reproduce locally| RCAKIND
RCACI --> ROOT[Root Cause Found]
RCAHS --> ROOT
RCAKIND --> ROOT
ROOT --> TDD["tdd:*"]:::tdd
classDef rca fill:#FF5722,stroke:#333,color:white
classDef tdd fill:#4CAF50,stroke:#333,color:white
Follow this diagram as the workflow.
RCA Skills
Root cause analysis workflows for systematic failure investigation.
Auto-Select Sub-Skill
When this skill is invoked, determine the right sub-skill based on context:
Step 1: Determine what's available
Check for HyperShift cluster:
ls ~/clusters/hcp/kagenti-hypershift-custom-*/auth/kubeconfig 2>/dev/null
Check for Kind cluster:
kind get clusters 2>/dev/null
Step 2: Route based on failure source and access
Where did the failure occur?
โ
โโ CI pipeline (GitHub Actions) โโโโโโโโโโโโโโโโโโโโโโโโโโ
โ โ
โ Do you have a live cluster matching the CI env? โ
โ โ โ
โ โโ HyperShift cluster available โ
โ โ โ Use `rca:hypershift` (deep investigation) โ
โ โ โ
โ โโ Kind cluster available (for Kind CI failures) โ
โ โ โ Use `rca:kind` (reproduce locally) โ
โ โ โ
โ โโ No cluster โ
โ โ Use `rca:ci` (logs and artifacts only) โ
โ โ If inconclusive, ask user to create cluster โ
โ โ
โโ Local Kind cluster โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ Use `rca:kind` (full local access) โ โ
โ โ โ
โโ HyperShift cluster โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ Use `rca:hypershift` (full remote access) โโ โ
โโ โ
After RCA is complete, switch to TDD for fix iteration: โโโโโ โ
- `tdd:ci` (CI-only) โ
- `tdd:hypershift` (live cluster) โ
- `tdd:kind` (local cluster) โ
Available Skills
| Skill | Access | Auto-approve | Best for |
|---|
rca:ci | CI logs/artifacts only | N/A | CI failures, no cluster |
rca:hypershift | Full cluster access | All read ops | Deep investigation |
rca:kind | Full local access | All ops | Kind failures, fast repro |
Concurrency limit: Only one rca:kind session at a time (one Kind cluster fits locally).
Before routing to rca:kind, run kind get clusters โ if a cluster exists from another session,
route to rca:ci instead or ask the user.
Related Skills
tdd:ci - Fix iteration after RCA (CI-driven)
tdd:hypershift - Fix iteration with live cluster
tdd:kind - Fix iteration on Kind
k8s:logs - Query and analyze component logs
k8s:pods - Debug pod issues