Skip to main content

yuguo-Jack/cuda-optimized-skill

SkillsMP は yuguo-Jack/cuda-optimized-skill から 4 件の skill を収集しています。skill を開くとソースと詳細を確認できます。

記録された最新のソース活動
SkillsMP カタログ更新
収集済み skills
4
GitHub スター
3
GitHub フォーク
0

このリポジトリの skills

1 件の職業カテゴリ · 100% 分類済み

収集済み skill 4 件中 4 件を表示しています。

職業分類
ソフトウェア開発者
説明

Capture, triage, benchmark, and optimize TorchInductor or hand-written Triton kernels on Hygon DCU gfx936/gfx938. Use when investigating torch.compile or TorchInductor Triton performance on DCU, parsing Triton autotune logs, saving generated kernels and…

原文の言語: 英語

更新
職業分類
ソフトウェア開発者
説明

Run remote compile, test, profiling, and debug tasks through SSH plus docker exec while keeping code edits local and synced to the remote node. Use when Codex must validate environment readiness, check ROCm/DTK/Hygon GPU card status, inspect Python packages,…

原文の言語: 英語

更新
職業分類
ソフトウェア開発者
説明

Iteratively optimize Hygon DCU HIP / CK Tile kernels against a Python reference using hipprof, DTK tools, dccobjdump ISA verification, roofline-style budgeting, branch selection, ablation attribution, and gfx936/gfx938-aware optimization references. Use when…

原文の言語: 英語

更新
職業分類
ソフトウェア開発者
説明

Generate a Hygon DCU HIP/C++ baseline kernel and correctness harness from a Torch, Triton, TileLang, Python, or CUDA/C++ reference plus shape JSON, including evidence-backed CUDA-to-HIP/DCU conversion, then hand the validated baseline to the Hygon HIP kernel…

原文の言語: 英語

更新
収集済み skill 4 件中 4 件を表示しています。