Skip to main content

vllm

Comprehensive reference documentation and skill for vLLM - a high-throughput and memory-efficient inference and serving engine for large language models (LLMs). Covers vLLM architecture (V0 and V1), engine APIs (LLMEngine, AsyncLLMEngine, LLM), OpenAI-compatible API server, configuration system, model executor and layers, attention mechanisms (PagedAttention, FlashAttention, MLA), KV cache management, sampling and decoding (beam search, speculative decoding, structured outputs), distributed inference (tensor/pipeline/data/expert parallelism), multimodal processing (image/audio/video), compilation and CUDA graphs, custom kernels, quantization (FP8, GPTQ, AWQ, INT4/INT8, etc.), LoRA and adapters, scheduling and memory management, supported model architectures (200+ models), speculative decoding (EAGLE, Medusa, n-gram), observability and profiling, and hardware platforms (NVIDIA GPU, AMD GPU, CPU, TPU, Intel GPU/XPU).

Zur Installation springen

Quellinformationen

Repository
jstzwj/ai-infra-plugins
Letzte Quellaktivität
7. Mai 2026 um 11:21
Erkannte Sprache von SKILL.md
Englisch
Sterne
4
Forks
0

Installationsoptionen

Standardmäßig ist der Prompt ausgewählt, der zuerst die Quelle prüft. Sie können zu einem direkten Befehl wechseln oder eine lokale Kopie herunterladen.

Quelldateien prüfen

Lesen Sie SKILL.md und alle von SkillsMP angezeigten Begleitdateien, bevor Sie sich für eine Installation entscheiden.