Skip to main content

yuguo-Jack/cuda-optimized-skill

SkillsMP 已收集 yuguo-Jack/cuda-optimized-skill 中的 4 个 Skill。打开任一 Skill 可查看来源和详情。

最近记录的来源活动
SkillsMP 收录数据更新
已收集 skills
4
GitHub 星标
3
GitHub Forks
0

这个仓库中的 skills

1 个职业分类 · 已分类 100%

已展示 4 / 4 个已收集 Skill。

职业分类
软件开发工程师
描述

Capture, triage, benchmark, and optimize TorchInductor or hand-written Triton kernels on Hygon DCU gfx936/gfx938. Use when investigating torch.compile or TorchInductor Triton performance on DCU, parsing Triton autotune logs, saving generated kernels and…

原文语言:英语

更新
职业分类
软件开发工程师
描述

Run remote compile, test, profiling, and debug tasks through SSH plus docker exec while keeping code edits local and synced to the remote node. Use when Codex must validate environment readiness, check ROCm/DTK/Hygon GPU card status, inspect Python packages,…

原文语言:英语

更新
职业分类
软件开发工程师
描述

Iteratively optimize Hygon DCU HIP / CK Tile kernels against a Python reference using hipprof, DTK tools, dccobjdump ISA verification, roofline-style budgeting, branch selection, ablation attribution, and gfx936/gfx938-aware optimization references. Use when…

原文语言:英语

更新
职业分类
软件开发工程师
描述

Generate a Hygon DCU HIP/C++ baseline kernel and correctness harness from a Torch, Triton, TileLang, Python, or CUDA/C++ reference plus shape JSON, including evidence-backed CUDA-to-HIP/DCU conversion, then hand the validated baseline to the Hygon HIP kernel…

原文语言:英语

更新
已展示 4 / 4 个已收集 Skill。