Skip to main content
Manusで任意のスキルを実行
ワンクリックで

cutedsl-kernels

スター24
フォーク2
更新日2026年7月7日 11:27

Write, debug, validate, and optimize NVIDIA CuTe DSL (nvidia-cutlass-dsl) GPU kernels in Python. Use for CuTe layout algebra, TV layouts, tiled copies, predication, shared/register/tensor memory, cp.async and TMA pipelines, mbarriers, warp specialization, warp/block/cluster reductions, MMA on SM80 (tensor cores), SM90 (WGMMA), and SM100/Blackwell (tcgen05/UMMA, TMEM, block-scaled FP4/FP8), tile schedulers, torch interop via from_dlpack, PTX inspection, and diagnosing alignment/vectorization/synchronization failures.

インストール

Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。

SKILL.md
readonly