Skip to main content
تشغيل أي مهارة في Manus
بنقرة واحدة

gpu-incident-triage

النجوم٠
التفرعات٠
آخر تحديث٢٣ يوليو ٢٠٢٦ في ٠٣:١٢

Triage a GPU HARDWARE incident on a K8s GPU node (Xid errors, NVLink faults, UnexpectedAdmissionError, allocatable drop, GPU fell off the bus): collect a deterministic evidence bundle (kernel NVRM/Xid via node debug pod, device-plugin and DCGM logs, driver-pod nvidia-smi, node state, admission events, GPU pod timeline), apply safe first-response (cordon), upload the bundle to Google Drive, and post a detailed issue to Slack #problem-of-thakicloud via qa-issue-dedup-poster with thread follow-ups. Use when "GPU 장애", "Xid 에러", "NVLink 오류", "GPU가 안 잡혀요", "UnexpectedAdmissionError", "allocatable 줄었다", "GPU requires reset", "노드 GPU 죽음", "gpu incident", "hardware fault triage". Do NOT use for GPU utilization/capacity questions (use h200-gpu-usage-inspector), app-level demo QA (use ai-platform-autoqa / demo-* RCA skills), or GPU purchase planning (use gpu-capacity-procurement-planner).

التثبيت

التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.

SKILL.md
readonly