Skip to main content
在 Manus 中运行任何 Skill
一键导入

llm-inference-serving

星标23
分支7
更新时间2026年5月7日 21:51

Deep LLM-inference-serving operational intuition — vLLM, SGLang RadixAttention, TensorRT-LLM tradeoffs; PagedAttention and KV-cache math, chunked prefill, prefix caching, disaggregated serving (DistServe, Mooncake), speculative decoding, quantization, MoE parallelism, TTFT/TPOT. Load when tuning serving-engine knobs, KV-cache sizing, batching design, quantization-format selection, speculative-decoding choice, or MoE layout. Skip for fine-tuning, embeddings serving, or hosted-API integration. Triggers on: "vLLM", "SGLang", "TensorRT-LLM", "PagedAttention", "RadixAttention", "chunked prefill", "prefix caching", "kv-cache-dtype", "EAGLE-3", "Medusa", "MLA", "AWQ", "NVFP4", "DistServe", "Mooncake", "EPLB", "TTFT", "TPOT", "goodput".

安装

用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。

SKILL.md
readonly