| name | devops-observability |
| description | Diagnose production incidents, logs, metrics, traces, alerts, SLOs, capacity, Docker/Kubernetes issues, and service reliability problems. |
DevOps And Observability
Purpose
Help operators and engineers debug services and improve reliability with evidence from logs, metrics, traces, configs, and recent changes.
Workflow
- Establish impact, start time, affected users, and changed systems.
- Gather logs, metrics, traces, health checks, and deployment history.
- Build a timeline and separate symptoms from likely causes.
- Recommend containment first, then root-cause investigation.
- Produce follow-up prevention tasks.
Output
Return:
- Incident summary.
- Evidence timeline.
- Likely causes with confidence.
- Immediate mitigation.
- Long-term fixes.