一键导入
linux-troubleshooting
Linux system troubleshooting workflow for diagnosing and resolving system issues, performance problems, and service failures.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Linux system troubleshooting workflow for diagnosing and resolving system issues, performance problems, and service failures.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Diagnose a broken/slow/crashlooping app in kubenuc or k8s-vms-daniele — gather evidence via Graylog/Grafana/kubectl before asking the user anything. Triggers on "X is down", "X is crashlooping", "X is erroring", "X is slow", "why did X restart".
Route documentation/notes to the right destination (Confluence, README/CLAUDE.md, Obsidian, or NetBox) based on content type, and confirm the exact destination with the user before writing anything. Also invoke proactively — before opening a PR — when the change is a new operational gotcha, an incident, a security-hardening trade-off, a version-pin rationale, a new app/cluster, or a structural Terraform/infra change, to decide whether README.md, a CLAUDE.md, or Confluence needs updating.
Cluster lifecycle operations — setting up new clusters, kubenuc-test overlay pattern, kustomize patch conventions, and FluxCD bootstrap structure for infra-cd clusters.
Before adding any new Prometheus scrape target or exporter, check the active-series budget in Grafana Cloud and propose write_relabel drops if needed. Triggers on "add a scrape config", "add an exporter", "enable metrics for X", "add a ServiceMonitor". Gates any change that could grow active series count.
Review a Renovate PR before merging — read the changelog, diff values against this repo's overrides, check e2e status, flag major bumps, and produce a merge checklist. Triggers on "review this renovate PR", "should I merge this chart bump", "is this renovate PR safe".
Ansible playbook/role conventions, Molecule testing patterns, and ansible-agent invocation for the infra-cd repository.
| name | linux-troubleshooting |
| description | Linux system troubleshooting workflow for diagnosing and resolving system issues, performance problems, and service failures. |
| category | granular-workflow-bundle |
| risk | safe |
| source | personal |
| date_added | 2026-02-27 |
Specialized workflow for diagnosing and resolving Linux system issues including performance problems, service failures, network issues, and resource constraints.
Use this workflow when:
bash-linux - Linux commandsdevops-troubleshooter - Troubleshootinguptime
hostnamectl
cat /etc/os-release
dmesg | tail -50
Use @bash-linux to gather system information
bash-linux - Resource commandsperformance-engineer - Performance analysistop -bn1 | head -20
free -h
df -h
iostat -x 1 5
Use @performance-engineer to analyze system resources
bash-linux - Process commandsserver-management - Process managementps aux --sort=-%cpu | head -10
pstree -p
lsof -p PID
strace -p PID
Use @server-management to investigate processes
bash-linux - Log commandserror-detective - Error detectionjournalctl -xe
tail -f /var/log/syslog
grep -i error /var/log/*
Use @error-detective to analyze log files
bash-linux - Network commandsnetwork-engineer - Network troubleshootingip addr show
ss -tulpn
curl -v http://target
dig domain
Use @network-engineer to diagnose network issues
server-management - Service managementsystematic-debugging - Debuggingsystemctl status service
journalctl -u service -f
systemctl restart service
Use @systematic-debugging to troubleshoot service issues
incident-responder - Incident responsebash-pro - Fix implementationUse @incident-responder to implement resolution
os-scripting - OS scriptingbash-scripting - Bash scriptingcloud-devops - DevOps