| name | kernel-perf-testing |
| description | Run TLX kernel performance benchmarks on Hopper, Blackwell, and AMD (gfx950/CDNA4, gfx1250) GPUs. Use when user asks to benchmark, profile, or measure performance of any TLX kernel (GEMM, Flash Attention, addmm+GLU, IKBO variants). Handles GPU selection, denoise wrapping (NVIDIA only), and version flags. Never run unless explicitly asked.
|
| disable-model-invocation | true |
Kernel Performance Testing
Never run performance tests unless the user explicitly asks.
Perf scripts live in third_party/tlx/tutorials/testing/test_<arch>_<op>_perf.py,
one per (op, input-contract, arch). Each takes [--version ...] to select a
kernel variant; with no --version it runs all variants for that op.
NVIDIA (Hopper / Blackwell)
GPU selection protocol
- Run
nvidia-smi to check GPU occupancy.
- Pick the GPU with the lowest memory usage.
- Set
CUDA_VISIBLE_DEVICES to that GPU.
Wrap benchmarks with denoise.sh (locks clocks/power) for stable results.
Hopper GPU
CUDA_VISIBLE_DEVICES=<gpu_id> third_party/tlx/denoise.sh python third_party/tlx/tutorials/testing/test_hopper_gemm_perf.py [--version {ws|pipelined}]
CUDA_VISIBLE_DEVICES=<gpu_id> third_party/tlx/denoise.sh python third_party/tlx/tutorials/testing/test_hopper_fa_perf.py [--version {ws|ws_pipelined|ws_pipelined_pingpong|ws_pipelined_pingpong_persistent}]
Blackwell GPU
CUDA_VISIBLE_DEVICES=<gpu_id> third_party/tlx/denoise.sh python third_party/tlx/tutorials/testing/test_blackwell_gemm_perf.py [--version {ws|clc|pipelined|2cta}]
CUDA_VISIBLE_DEVICES=<gpu_id> third_party/tlx/denoise.sh python third_party/tlx/tutorials/testing/test_blackwell_fa_perf.py [--version {ws|ws_persistent|ws_pipelined|ws_pipelined_persistent|clc}] [--mode fwd|bwd]
CUDA_VISIBLE_DEVICES=<gpu_id> third_party/tlx/denoise.sh python third_party/tlx/tutorials/testing/test_blackwell_fa_mxfp8_perf.py