Skip to main content
Run any Skill in Manus
with one click

pytorch-perf-training

Stars17
Forks23
UpdatedMay 19, 2026 at 06:50

Optimise PyTorch single-GPU training wall-clock time. torch.compile, CUDA graphs, Tensor Cores, AMP, profiling, memory pre-allocation, Triton kernels, compiled optimizers, gradient monitoring, model visualization, benchmarking. Use when training DiT, transformers, diffusion models on GPU and user mentions speed, throughput, compilation, graph breaks, reduce-overhead, profiling, nsight, kernel fusion, gradient flow, vanishing/exploding gradients, hooks, torchinfo, torchviz, record_function, benchmark, or training step timing. Also trigger for looped/recursive transformers with shared weights, weight tying, depth-wise LoRA, or routing in looped blocks.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

File Explorer
12 files
SKILL.md
readonly