Skip to main content
Run any Skill in Manus
with one click

tritonify

Stars4
Forks2
UpdatedJune 26, 2026 at 08:05

Agent-driven Triton/CUDA kernel optimization: a roofline-targeted trial-loop that treats cuBLAS/cuDNN/Liger as baselines to BEAT — never claiming an unmeasured speedup, never calling an op impossible without checking GPU access. Use to write, optimize, fuse, profile, port, or speed up any Triton/CUDA kernel or LLM op — GEMM, MLP, MoE, attention, activation, fused/custom loss, quantized.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

File Explorer
9 files
SKILL.md
readonly