一键导入
pytorch-nan-debugging
Use when debugging NaN/inf instability, exploding losses, or numerical collapse in PyTorch training runs.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Use when debugging NaN/inf instability, exploding losses, or numerical collapse in PyTorch training runs.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Use when opening a PR and driving it to green CI — push the branch, write a high-level PR description, watch checks, fix failures, and keep looping until green; also use when an already-open PR gets new commits.
Use when the user starts, resumes, switches, saves, or recalls a named task or ongoing work across sessions (e.g. "remember this", "continue X", "what was I doing on Y", "track this task").
Git commit workflow. Load when finishing any code-change task (after verification) to commit and push, or when explicitly staging/committing.
Use when modifying or adding a transform, loss, model, method, package, task-model, config, docs page, or example in a Lightly AG repo (lightly-ssl or lightly-train) and you need repo-specific conventions.
Spawn a subagent to run a task. Use when you want to delegate work to a separate pi instance.
Use when configuring or replicating LT-DETR v2 (ltdetrv2-s/m/l/x) COCO benchmarks via the lightly-train wrapper — model alias map, global vs per-rank batch size, recipe invariants, ECDet config mapping, cluster pins, and common gotchas.
| name | pytorch-nan-debugging |
| description | Use when debugging NaN/inf instability, exploding losses, or numerical collapse in PyTorch training runs. |
Use this workflow to localize the first non-finite value, not just the final failing op.
torch.autograd.set_detect_anomaly(True).TORCH_SHOW_CPP_STACKTRACES=1 and CUDA_LAUNCH_BLOCKING=1.<out>/checkpoints/last.ckpt, copy the source checkpoint there first.detect_anomaly identifies the failing autograd op, not always the exact module.Report: