| name | gpu-incident-triage |
| description | Triage a GPU HARDWARE incident on a K8s GPU node (Xid errors, NVLink faults, UnexpectedAdmissionError, allocatable drop, GPU fell off the bus): collect a deterministic evidence bundle (kernel NVRM/Xid via node debug pod, device-plugin and DCGM logs, driver-pod nvidia-smi, node state, admission events, GPU pod timeline), apply safe first-response (cordon), upload the bundle to Google Drive, and post a detailed issue to Slack #problem-of-thakicloud via qa-issue-dedup-poster with thread follow-ups. Use when "GPU ์ฅ์ ", "Xid ์๋ฌ", "NVLink ์ค๋ฅ", "GPU๊ฐ ์ ์กํ์", "UnexpectedAdmissionError", "allocatable ์ค์๋ค", "GPU requires reset", "๋
ธ๋ GPU ์ฃฝ์", "gpu incident", "hardware fault triage". Do NOT use for GPU utilization/capacity questions (use h200-gpu-usage-inspector), app-level demo QA (use ai-platform-autoqa / demo-* RCA skills), or GPU purchase planning (use gpu-capacity-procurement-planner). |
GPU Incident Triage (ํ๋์จ์ด ์ฌ๊ณ ์ด๋ ๋์)
2026-07-23 gpu-003 ์ฌ๊ณ (Xid 74 NVLink fatal 3-GPU ๋์ โ Xid 79 bus ์ดํ โ ์ GPU ERR!)์์
๊ฒ์ฆ๋ ์ ์ฐจ๋ฅผ ์คํฌํ. ์์งยท๋ ๋๋ ์ฝ๋ ์์ ([[sonnet-format-determinism]]), ๊ฒ์๋ ์น์ธ ๊ฒ์ดํธ,
ํ๊ดด์ ์กฐ์น(๋ฆฌ๋ถํธยทํ ํ ํฌ๋ drain)๋ ํ์ง ์๋๋ค โ ์ฆ๊ฑฐ๋ฅผ ๋ง๋ค์ด ๊ด๋ฆฌ์์๊ฒ ๋๊ธด๋ค.
์ ์ฐจ (์์ ๊ณ ์ )
1. ์์ง (๊ฒฐ์ ๋ก ์์ง๊ธฐ โ ๋ฌด๋ณ๊ฒฝ read-only)
bash scripts/skills/gpu_incident_collect.sh <kube-context> <node> [tag]
์ฐ์ถ: outputs/incidents/<date>-<node>-<tag>/ + zip. ํฌํจ: ์ปค๋ NVRM/Xid(journalctl, debug pod),
device-plugin/DCGM ๋ก๊ทธ, driver-pod nvidia-smi(-q ํฌํจ), node describe/yaml, admission ์ด๋ฒคํธ,
GPU ํฌ๋ ํ์๋ผ์ธ. INCIDENT-SUMMARY.md๋ฅผ ์์ผ๋ก ์ถ๊ฐ(ํ์๋ผ์ธยทํ์ ยท์กฐ์น ์ด๋ ฅ โ ์๋ 4์ ๋ด์ฉ).
ํต์ฌ ํ๋
๋ฒ:
Xid 79 = GPU fell off the bus (PCIe) โ ๊ฐ๋ณ ๋ฆฌ์
๊ฑฐ์ ๋ถ๊ฐ, ๋ฆฌ๋ถํธ ํ์ค.
Xid 74 = NVLink fatal. ์ฌ๋ฌ GPUยท์ฌ๋ฌ ๋งํฌ ๋์๋ฉด ์นด๋ ๋ถ๋์ด ์๋๋ผ ๋ธ๋ฆฌ์ง/ํจ๋ธ๋ฆญ ์ด๋ฒคํธ.
- device-plugin์
XidCriticalError ... marking device as unhealthy โ allocatable ๊ฐ๋ฑ์ ์ง์ ์์ธ.
- nvidia-smi์์
Unable to determine the device handle / ERR! / ํ๋ก์ธ์ค ์๋ ๋ฉ๋ชจ๋ฆฌ ์ข์ด ํ์ธ.
- ์ฌ๊ณ ์๊ฐ ์ ํ GPU ํฌ๋ ํ์๋ผ์ธ์ผ๋ก ๋๊ฐ ๊ทธ ์๊ฐ๋์ ์ผ๋์ง ์ ์งํ๊ฒ ๊ธฐ๋ก(์๊ดโ ์ธ๊ณผ ๋ช
์).
2. ์์ ์ด๋ ์กฐ์น (๊ฐ์ญ๋ง)
kubectl --context <ctx> cordon <node>
- โ ํ ํ Running ํฌ๋ drain ๊ธ์ง(ํ์ ์ฌํญ์ผ๋ก ์ด์์ ๊ธฐ์ฌ). โ ๋
ธ๋ ๋ฆฌ๋ถํธ ์ง์ ์ํ ๊ธ์ง(๊ด๋ฆฌ์).
- GPU ๋ฆฌ์
์ ์๋ ์์ฒด๊ฐ ์ฆ๊ฑฐ: NVLink ๋๋ฉ์ธ์ ํ์ด ์ ์ฒด ๋ฆฌ์
์๊ตฌ + off-bus GPU๋ ํธ๋ค ์์
โ ๊ฑฐ๋ถ ๋ฉ์์ง๋ฅผ ๋ฒ๋ค์ ๋จ๊ธฐ๊ณ "๋ฆฌ๋ถํธ ํ์"์ ๊ทผ๊ฑฐ๋ก ์ด๋ค.
3. Drive ์
๋ก๋ (์ด๋ ๊ณต๊ฐ)
gws drive +upload "<abs>/<bundle>.zip" --format json
gws drive permissions create --params '{"fileId":"<FID>"}' --json '{"role":"reader","type":"anyone"}'
4. Slack ๊ฒ์ (qa-issue-dedup-poster ๊ฒฝ์ โ ์ฑ๋ ์ค๋ฒ๋ผ์ด๋)
Skill: qa-issue-dedup-poster ์ํฌํ๋ก๋ฅผ ๋ฐ๋ฅด๋ --channel C0A9AUPQ62X(#problem-of-thakicloud):
sync โ dedup โ post.py ๋ ๋(์กด๋๋ง ํฉ๋๋ค์ฒด, ํ์ /[์ถ์ ] ๊ตฌ๋ถ, ํ
๋ง ๋ถ๋ฆฟ) โ --post.
--domain ์ธํ๋ผ(๋ฐ์ค์ฌ ์๋ ํ๊ทธ) + ํ๋์จ์ด ๋ด๋น --assignee(์: ๊น์ข
์ U0A6KDMSTB5).
- ์ค๋ ๋ 2๊ฑด ํ์: (1) Drive ๋ฒ๋ค ๋งํฌ + ํต์ฌ ์์ฝ(Xid ํ์๋ผ์ธยทํ ์ํยท๋ฒ๋ค ๊ตฌ์ฑ),
(2) ์กฐ์น ์ํฉ(์์ง/cordon/๋ฆฌ์
๋ถ๊ฐ ๊ทผ๊ฑฐ/diag ๋ณด๋ฅ โ ๊ฐ ์๋ฃยท๋ถ๊ฐยท๋ณด๋ฅ ๋ช
์).
- ๊ฒ์๋ ์ฌ์ฉ์ ์น์ธ ํ์(Safety Gates). ๋ ๋ ๋๋ผ์ด๋ฐ์ ๋จผ์ ๋ณด์ฌ์ค๋ค.
5. ๋ง๋ฌด๋ฆฌ
- ํ/๋
ธํธ์ ์ฌ๊ณ ๊ธฐ๋ก([[aps-failure-archaeology]] ์ธ๋ฑ์ค ๋์). ๋ณต๊ตฌ ํ์ธ ํ
uncordon + ์ค๋ ๋์ ์ข
๊ฒฐ ์ฝ๋ฉํธ.
- ๋ฆฌ๋ถํธ ํ ๊ถ์ฅ:
dcgmi diag -r 3, Xid 74 ์ฌ๋ฐ ์ NVLink ๋ธ๋ฆฌ์ง ๋ฆฌ์ํธ/์ผ์ด๋ธ โ RMA ๊ฒฝ๋ก.
gotchas (2026-07-23 ์ค์ฌ๊ณ ์์)
- busybox
dmesg๋ ๊ถํ ์์ โ chroot /host journalctl -k ๊ฐ ์ ๋ต (--profile=sysadmin ํ์).
- ํธ์คํธ PATH์ nvidia-smi ์์ ์ ์์(
/run/nvidia/driver/usr/bin/) โ driver daemonset ํฌ๋ exec๊ฐ ์์ .
- UnexpectedAdmissionError ํฌ๋์๋ฃจํ๋ Deployment๊ฐ ์คํจ ํฌ๋๋ฅผ ์๋ฐฑ ๊ฐ ์์ฐํ๋ค โ
์์ธ ์ํฌ๋ก๋๊ฐ ์ฐ๋ฆฌ ๊ฒ์ด๋ฉด deployment ์ฆ์ ์ญ์ +
delete pods --field-selector=status.phase=Failed.
- ์ฌ๊ณ ๋
ธ๋์ k8s์ Running ํฌ๋๋ GPU ๋ ๋ฒจ์์ ์ฃฝ์ด ์์ ์ ์๋ค(ํ๋ก์ธ์ค ์๋ ๋ฉ๋ชจ๋ฆฌ ์ข์ด) โ
"Running์ด๋ ๊ด์ฐฎ๋ค"๊ณ ํ๋จ ๊ธ์ง, nvidia-smi๊ฐ ์ ๋ณธ.
- Slack user ํ ํฐ์ users.list/conversations.list ์ค์ฝํ๊ฐ ์๋ค โ ์ฑ๋/์ ์ ID๋ MCP
slack_search_channels/slack_search_users๋ก ์กฐํ.