Use when porting, tuning, or diagnosing TIRx CUDA kernels, especially for perf-gate failures, PTX/SASS divergence, bitwise mismatches, register spills, scoreboard stalls, address arithmetic, predication, uniformity, pipeline depth, TMA, shared-memory…
mlc-ai/TIRx-kernels
SkillsMP has collected 3 skills from mlc-ai/TIRx-kernels. Open a skill to review its source and details.
- Latest recorded source activity
- SkillsMP catalog refreshed
- skills collected
- 3
- GitHub stars
- 94
- GitHub forks
- 10
Skills in this repository
1 occupation categories · 33% classified
Showing 3 of 3 collected skills.
skill
occupation
description
updated
occupation
Software Developers
description
updated
occupation
unclassified
description
Integrate kernels into tirx-kernels using its current module, registry, licensing, correctness, benchmark, and bench-suite conventions. Use when adding, porting, testing, benchmarking, or reviewing a kernel in the repository.
updated
occupation
unclassified
description
Port performance kernels from CUDA, CuTeDSL, Gluon, and Triton to TIRx through five gated stages: scaffolding, kernel sketch, sketch reviewer, correctness gate, and performance gate. Use when the user asks to port or align an optimized kernel implementation…
updated
Showing 3 of 3 collected skills.