Skip to main content
Run any Skill in Manus
with one click

cuda-dev-routing

Stars6
Forks0
UpdatedJuly 17, 2026 at 07:56

Run BEFORE writing a CUDA/Triton kernel, adopting a GPU library, or optimizing slow training/benchmark/eval throughput — routes the work to the lever that actually pays on this fleet. Fires on "speed up training", "the training is slow", "write a CUDA kernel", "Triton kernel", "GPU-accelerate this", "make the eval faster", "optimize throughput", "use CUDA", "fuse this op", "should this run on the GPU", "the benchmark takes forever", or any plan whose first move is to write a kernel. Also fires when a session proposes FP8/MXFP4 training on the Spark, flash-attn on sm_121, GPU vector search / GPU dedup for the RAG corpus, Muon for fine-tuning, or Unsloth — each is a REFUTED lever with a specific refuting fact below, and re-attempting one is the exact repeat-waste this skill exists to stop. Applies to the DGX Spark (GB10/sm_121/aarch64), the RTX PRO 6000 (sm_120/x86), RunPod pods, and the Mac MLX boxes.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

SKILL.md
readonly