Skip to main content
Run any Skill in Manus
with one click

cutedsl-kernels

Stars24
Forks2
UpdatedJuly 7, 2026 at 11:27

Write, debug, validate, and optimize NVIDIA CuTe DSL (nvidia-cutlass-dsl) GPU kernels in Python. Use for CuTe layout algebra, TV layouts, tiled copies, predication, shared/register/tensor memory, cp.async and TMA pipelines, mbarriers, warp specialization, warp/block/cluster reductions, MMA on SM80 (tensor cores), SM90 (WGMMA), and SM100/Blackwell (tcgen05/UMMA, TMEM, block-scaled FP4/FP8), tile schedulers, torch interop via from_dlpack, PTX inspection, and diagnosing alignment/vectorization/synchronization failures.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

SKILL.md
readonly