benchmark-harness
Use when creating or revising CUDA benchmark runners and result artifacts.
Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.
Menu
Use when creating or revising CUDA benchmark runners and result artifacts.
Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.
Based on SOC occupation classification
Use when CUDA availability, compiler setup, or profiler visibility may be broken or inconsistent.
Use when designing or reviewing handwritten CUDA FlashAttention kernels.
Use when designing or reviewing handwritten CUDA GEMM kernels.
Use for kernel-level analysis with Nsight Compute after a benchmark is reproducible.
Use for stream overlap, launch overhead, and end-to-end CUDA timeline analysis with Nsight Systems.
Use when deciding whether a CUDA kernel is likely memory-bound or compute-bound.
| name | benchmark-harness |
| description | Use when creating or revising CUDA benchmark runners and result artifacts. |
Use this skill when creating or updating benchmark runners.
artifacts/benchmarks/.