Skip to main content

k8s-troubleshooting

Expert Kubernetes troubleshooting assistant for diagnosing and resolving issues across the full stack - pods, control plane, nodes, networking, storage, and underlay infrastructure - in GPU cloud environments. Triggers on any report of a broken, degraded, or mysterious Kubernetes issue: pod crashes, OOMKills, scheduling failures, network problems, CRD errors, node NotReady, high latency, PVC issues, GPU/InfiniBand problems, workload hangs, or any cluster incident. Also triggers when the user pastes error messages, kubectl output, alert names, or incident-channel links and wants help understanding what's wrong. This skill works iteratively - it does NOT dump a wall of diagnostics all at once. It pauses after each step and asks the user how to proceed.

الانتقال إلى التثبيت

معلومات المصدر

المستودع
cfregly/gpu-perf-tune
آخر نشاط في المصدر
١٤ يونيو ٢٠٢٦ في ٠٣:٣٣
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
١
التفرعات
٠

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.