Skip to main content

mlc-ai/TIRx-kernels

SkillsMP 已收集 mlc-ai/TIRx-kernels 中的 3 个 Skill。打开任一 Skill 可查看来源和详情。

最近记录的来源活动
SkillsMP 收录数据更新
已收集 skills
3
GitHub 星标
94
GitHub Forks
10

这个仓库中的 skills

1 个职业分类 · 已分类 33%

已展示 3 / 3 个已收集 Skill。

职业分类
软件开发工程师
描述

Use when porting, tuning, or diagnosing TIRx CUDA kernels, especially for perf-gate failures, PTX/SASS divergence, bitwise mismatches, register spills, scoreboard stalls, address arithmetic, predication, uniformity, pipeline depth, TMA, shared-memory…

原文语言:英语

更新
职业分类
未分类
描述

Integrate kernels into tirx-kernels using its current module, registry, licensing, correctness, benchmark, and bench-suite conventions. Use when adding, porting, testing, benchmarking, or reviewing a kernel in the repository.

原文语言:英语

更新
职业分类
未分类
描述

Port performance kernels from CUDA, CuTeDSL, Gluon, and Triton to TIRx through five gated stages: scaffolding, kernel sketch, sketch reviewer, correctness gate, and performance gate. Use when the user asks to port or align an optimized kernel implementation…

原文语言:英语

更新
已展示 3 / 3 个已收集 Skill。