원클릭으로
llm-kernel-map
Map complete kernel pipeline for LLM inference (prefill vs decode)
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
메뉴
Map complete kernel pipeline for LLM inference (prefill vs decode)
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
SOC 직업 분류 기준
Profile LLM with ATOM, extract GPU kernels, generate kernel inventory
Extract per-layer kernel sequence from ATOM level=3 traces
Launch ATOM server and run throughput benchmark
CUDA graph replay benchmark + source trace templates
Single GPU kernel micro-benchmark with roofline analysis
Deep kernel source-level walkthrough for aiter/atom/CK stack
| name | llm-kernel-map |
| description | Map complete kernel pipeline for LLM inference (prefill vs decode) |
| user-invocable | true |
Maps out the complete kernel pipeline for LLM inference, identifying prefill vs decode differences.
User asks about what kernels an LLM uses, how inference works end-to-end, or wants to compare different model architectures at the kernel level.
tokens → Embedding → [Decoder Layer × N] → Final Norm → LM Head → Sampling → token_id
Prefill (multi-token): Decode (single-token):
RMSNorm (pre-attn) RMSNorm (pre-attn)
QKV GEMM (large M) QKV GEMM (skinny M)
RoPE RoPE
KV Cache Store KV Cache Store
Flash Attention (fmha) Paged Attention (PA) + Reduce
O GEMM (large M) O GEMM (skinny M)
RMSNorm + Residual Add RMSNorm + Residual Add
Gate+Up GEMM (large M) Gate+Up GEMM (skinny M)
SiLU activation SiLU activation
Down GEMM (large M) Down GEMM (skinny M)
RMSNorm + Residual Add RMSNorm + Residual Add
Same as above but MLP replaced by:
Gate GEMM → TopK Softmax → Token Sort → Expert GEMM (2-stage) → Accumulate
Same as above but Attention replaced by:
Down-proj Q/KV → Latent attention → Up-proj → (no explicit KV cache separation)
| Dimension | Prefill | Decode |
|---|---|---|
| GEMM shape | M=seq_len (large) | M=1 (skinny) |
| GEMM kernel | Large tile (MT64x32x256) | Skinny tile (MT32x16x256) |
| Attention | Flash Attention (O(N^2d)) | Paged Attention (O(Nd)) |
| MoE Sort | Multi-phase (many tokens) | Single-phase (1 token) |
| Bottleneck | Compute (GEMM, MoE) | Memory BW (PA reads KV cache) |
| Latency | Higher total, amortized per token | Lower per token |
| Function | Provider | Note |
|---|---|---|
| Embedding | PyTorch native / Triton | aten::index_select or triton_poi_fused_embedding |
| RMSNorm | aiter | Fused add+norm+quant, HIP C++ |
| RoPE | aiter | 2-channel inplace, HIP C++ |
| KV Cache | aiter | Per-token FP8 quant, HIP C++ |
| Flash Attention | aiter | Hand-written ASM (.co binary) |
| Paged Attention | aiter | Hand-written ASM (.co binary) |
| Dense GEMM | rocBLAS | Cijk_* kernels, auto-tuned tiles |
| MoE Gate | vllm | topkGatingSoftmax CUDA kernel |
| MoE Sort | CK (ck_tile) | MoeSortingKernel |
| MoE Expert GEMM | CK | kernel_moe_gemm (2-stage fused) |
| Sampling | aiter | Gumbel-max trick, HIP C++ |
Category Time% Bound-by
MoE Expert GEMM 31% Compute (prefill) / Memory (decode)
Dense GEMM (QKV/O) 23% Compute (prefill) / Memory (decode)
RMSNorm 12% Launch overhead (small n)
Paged Attention 10% Memory BW (reads KV cache)
MoE Gate+Sort 13% Launch overhead
RoPE + KV Cache 10% Launch overhead
Other 1% -
| Model | hidden | layers | heads | KV heads | head_dim | experts | topk |
|---|---|---|---|---|---|---|---|
| Qwen3-30B-A3B | 2048 | 48 | 32 | 4 | 128 | 128 | 8 |
| Qwen3-235B-A22B | 4096 | 94 | 64 | 4 | 128 | 128 | 8 |
| DeepSeek-V3 | 7168 | 61 | 128 | 128(MLA) | 128 | 256 | 8 |
| Llama-3-70B | 8192 | 80 | 64 | 8 | 128 | - | - |