Skip to main content
Run any Skill in Manus
with one click

llm-inference-serving

Stars23
Forks7
UpdatedMay 7, 2026 at 21:51

Deep LLM-inference-serving operational intuition — vLLM, SGLang RadixAttention, TensorRT-LLM tradeoffs; PagedAttention and KV-cache math, chunked prefill, prefix caching, disaggregated serving (DistServe, Mooncake), speculative decoding, quantization, MoE parallelism, TTFT/TPOT. Load when tuning serving-engine knobs, KV-cache sizing, batching design, quantization-format selection, speculative-decoding choice, or MoE layout. Skip for fine-tuning, embeddings serving, or hosted-API integration. Triggers on: "vLLM", "SGLang", "TensorRT-LLM", "PagedAttention", "RadixAttention", "chunked prefill", "prefix caching", "kv-cache-dtype", "EAGLE-3", "Medusa", "MLA", "AWQ", "NVFP4", "DistServe", "Mooncake", "EPLB", "TTFT", "TPOT", "goodput".

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

SKILL.md
readonly