Skip to main content
facebookexperimental
Perfil de criador do GitHub

facebookexperimental

Visão por repositório de 21 skills coletadas em 2 repositórios do GitHub.

skills coletadas
21
repositórios
2
atualizado
1 de set. de 2026
explorador de repositórios

Repositórios e skills representativas

trx-kernel-optimization-agent
sem classificação

Execute the TLX Kernel Optimization Agent CLI on a Triton or TLX kernel. Use this skill whenever the user says "use the TLX agent", "use the kernel optimization agent", "用 TLX agent 优化", or asks Carl/Claude to optimize a kernel with the repository agent. The…

1 de set. de 2026
tlx-api-reference
sem classificação

TLX DSL API reference for low-level GPU primitives. Use when writing or modifying TLX kernel code that uses barriers (mbarrier, named barriers), memory allocation (local_alloc, SMEM, TMEM), TMA operations, warp specialization (async_tasks, async_task), CLC…

20 de ago. de 2026
tlx-amd-testing
sem classificação

Test and run TLX-AMD tutorial kernels (gfx950/CDNA4 and gfx1250) and understand their CI. Use when working on AMD TLX tutorial kernels — GEMM (warp-pipeline, LDS-pipelined, TDM, MXFP), Flash Attention (simple, prefetch, persistent), addmm+GLU, or IKBO (FA,…

19 de ago. de 2026
running-with-buck
Desenvolvedores de software

How to build and run GPU targets under Buck in fbcode. Use when invoking buck2 run / buck2 build for any GPU benchmark, test, or kernel — selecting the GPU architecture and CUDA version, using @mode/opt and the beta Triton modifier, passing environment…

3 de ago. de 2026
sched2tlx-perf-testing
Analistas de garantia de qualidade de software e testadores

Run the sched2tlx perf/correctness harness over the modulo-scheduling example corpus (case1-9: GEMM, persistent GEMM, FA fwd/bwd, addmm+bias, LayerNorm, wgrad+bias, multiphase GEMM, scaled_mm). Use when the user asks to benchmark generated-vs-handwritten…

27 de jul. de 2026
barrier-visualization
Desenvolvedores de software

Produce a structured barrier report for AutoWS (automatic warp specialization) IR. Use when the user wants to visualize, audit, or debug barrier usage across warp-specialized partitions, or when debugging a GPU kernel hang (deadlock). For hangs, first dump IR…

1 de jul. de 2026
compute-sanitizer
Desenvolvedores de software

Run NVIDIA compute-sanitizer (memcheck, racecheck, initcheck, synccheck) against a Triton/TLX kernel to find runtime memory and synchronization bugs. Use when a kernel produces wrong results, crashes with an illegal/misaligned access, or is suspected of a…

26 de jun. de 2026
kernel-perf-testing
Desenvolvedores de software

Run TLX kernel performance benchmarks on Hopper, Blackwell, and AMD (gfx950/CDNA4, gfx1250) GPUs. Use when user asks to benchmark, profile, or measure performance of any TLX kernel (GEMM, Flash Attention, addmm+GLU, IKBO variants). Handles GPU selection,…

24 de jun. de 2026
Mostrando 8 de 16 skills coletadas.
Mostrando 2 de 2 repositórios
Todos os repositórios foram exibidos