Skip to main content
在 Manus 中运行任何 Skill
一键导入

pytorch-perf-training

星标17
分支23
更新时间2026年5月19日 06:50

Optimise PyTorch single-GPU training wall-clock time. torch.compile, CUDA graphs, Tensor Cores, AMP, profiling, memory pre-allocation, Triton kernels, compiled optimizers, gradient monitoring, model visualization, benchmarking. Use when training DiT, transformers, diffusion models on GPU and user mentions speed, throughput, compilation, graph breaks, reduce-overhead, profiling, nsight, kernel fusion, gradient flow, vanishing/exploding gradients, hooks, torchinfo, torchviz, record_function, benchmark, or training step timing. Also trigger for looped/recursive transformers with shared weights, weight tying, depth-wise LoRA, or routing in looped blocks.

安装

用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。

文件资源管理器
12 个文件
SKILL.md
readonly