Skip to main content
Jeden Skill in Manus ausführen
mit einem Klick

cuda-dev-routing

Sterne6
Forks0
Aktualisiert17. Juli 2026 um 07:56

Run BEFORE writing a CUDA/Triton kernel, adopting a GPU library, or optimizing slow training/benchmark/eval throughput — routes the work to the lever that actually pays on this fleet. Fires on "speed up training", "the training is slow", "write a CUDA kernel", "Triton kernel", "GPU-accelerate this", "make the eval faster", "optimize throughput", "use CUDA", "fuse this op", "should this run on the GPU", "the benchmark takes forever", or any plan whose first move is to write a kernel. Also fires when a session proposes FP8/MXFP4 training on the Spark, flash-attn on sm_121, GPU vector search / GPU dedup for the RAG corpus, Muon for fine-tuning, or Unsloth — each is a REFUTED lever with a specific refuting fact below, and re-attempting one is the exact repeat-waste this skill exists to stop. Applies to the DGX Spark (GB10/sm_121/aarch64), the RTX PRO 6000 (sm_120/x86), RunPod pods, and the Mac MLX boxes.

Installation

Mit Codex oder Claude installieren Kopieren Sie diesen Prompt, fügen Sie ihn in Codex, Claude oder einen anderen Assistant ein und lassen Sie die Skill-Seite prüfen und installieren.

SKILL.md
readonly