Skip to main content
在 Manus 中运行任何 Skill
一键导入

tsdb-diagnosis

// Diagnose training job incidents and check cluster health using the per-job Prometheus TSDB. Use when the user asks to diagnose a failure root cause, check GPU/network health, query Prometheus metrics, investigate a hang, or when the triage skill recommends deeper TSDB analysis.

$ git log --oneline --stat
stars:28
forks:3
updated:2026年5月3日 19:33
SKILL.md
readonly