benchmark-harness
星标15
分支0
更新时间2026年3月17日 09:40
Use when creating or revising CUDA benchmark runners and result artifacts.
安装
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
SKILL.md
readonly菜单
Use when creating or revising CUDA benchmark runners and result artifacts.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
Use when CUDA availability, compiler setup, or profiler visibility may be broken or inconsistent.
Use when designing or reviewing handwritten CUDA FlashAttention kernels.
Use when designing or reviewing handwritten CUDA GEMM kernels.
Use for kernel-level analysis with Nsight Compute after a benchmark is reproducible.
Use for stream overlap, launch overhead, and end-to-end CUDA timeline analysis with Nsight Systems.
Use when deciding whether a CUDA kernel is likely memory-bound or compute-bound.
基于 SOC 职业分类
| name | benchmark-harness |
| description | Use when creating or revising CUDA benchmark runners and result artifacts. |
Use this skill when creating or updating benchmark runners.
artifacts/benchmarks/.