Skip to main content

vllm-installer

This skill should be used when users need to install, configure, debug, or run vLLM inference server on NVIDIA GPUs (especially B200/H100/A100). It covers installation from PyPI or source, dependency management, environment setup, common error diagnosis and fixes, tensor parallelism configuration, and server startup/testing. The skill automatically checks for LSSD mount status and DeepEP installation for MoE models.

Quellinformationen

Repository
yangwhale/gpu-tpu-pedia
Letzte Quellaktivität
29. Januar 2026 um 10:42
Erkannte Sprache von SKILL.md
Englisch
Sterne
14
Forks
3

Installationsoptionen

Standardmäßig ist der Prompt ausgewählt, der zuerst die Quelle prüft. Sie können zu einem direkten Befehl wechseln oder eine lokale Kopie herunterladen.

Quelldateien prüfen

Lesen Sie SKILL.md und alle von SkillsMP angezeigten Begleitdateien, bevor Sie sich für eine Installation entscheiden.

Datei-Explorer
5 Dateien

SKILL.md wird angezeigt

SKILL.md
Quellanweisungen · Schreibgeschützte Vorschau
name
vllm-installer
description
This skill should be used when users need to install, configure, debug, or run vLLM inference server on NVIDIA GPUs (especially B200/H100/A100). It covers installation from PyPI or source, dependency management, environment setup, common error diagnosis and fixes, tensor parallelism configuration, and server startup/testing. The skill automatically checks for LSSD mount status and DeepEP installation for MoE models.
license
MIT
# vLLM Installer This skill provides comprehensive guidance for installing, configuring, and debugging vLLM on NVIDIA GPUs with CUDA 12.x. ## When to Use This Skill - Installing vLLM on NVIDIA GPUs (B200/H100/A100) - Debugging vLLM installation errors (missing libraries, version conflicts) - Configuring tensor parallelism for different model architectures - Setting up environment variables for CUDA and NVIDIA libraries - Starting and testing vLLM OpenAI-compatible API server - Fixing common runtime errors (cuDNN, cusparseLt, FlashInfer issues) ## Version Information (as of v0.14.1) | Component | Version | Notes | |-----------|---------|-------| | vLLM | 0.14.1 | Latest stable (v0.15.0rc2 not yet on PyPI) | | flashinfer-python | 0.5.3 | Attention backend | | flashinfer-cubin | 0.5.3 | Must match flashinfer-python version | | nixl | 0.9.0 | KV cache transfer (DMA-BUF, recommended for PD disaggregation) | | nvidia-nccl-cu12 | 2.28.3 | Force reinstall | | nvidia-cudnn-cu12 | 9.16.0.29 | Required for PyTorch 2.9+ | | bitsandbytes | 0.46.1 | Quantization support | | numpy | <2.3 | Required for numba compatibility | ### DeepSeek-V3 FP8 on Blackwell (B200) **vLLM correctly handles DeepSeek-V3 FP8 on Blackwell GPUs**, unlike SGLang which produces garbage output due to FP8 scale format incompatibility. | Framework | DeepSeek-V3 FP8 on Blackwell | |-----------|------------------------------| | **vLLM 0.14.1** | ✅ Works correctly | | SGLang 0.5.8 | ❌ Garbage output (scale format mismatch) | If you need to run DeepSeek-V3 on Blackwell (B200) GPUs, **use vLLM**: ```bash vllm serve deepseek-ai/DeepSeek-V3 \ --tensor-parallel-size 8 \ --port 8000 \ --download-dir /lssd/huggingface/hub \ --trust-remote-code \ --max-model-len 4096 ``` ## Pre-Installation Checks ### Step 0: Check Prerequisites Before installing vLLM, the skill automatically checks: 1. **LSSD Mount Status** - High-speed local SSD for model caching 2. **DeepEP Installation** - Required for MoE models (DeepSeek-V3, DeepSeek-R1) #### LSSD Check ```bash # Check if /lssd is mounted if mountpoint -q /lssd 2>/dev/null; then echo "✓ LSSD is mounted: $(df -h /lssd | tail -1 | awk '{print $2}')" else echo "✗ LSSD is not mounted" echo " Run: /lssd-mounter" fi ``` If LSSD is not mounted, use the `lssd-mounter` skill: ```bash /lssd-mounter ``` #### DeepEP Check (for MoE models) ```bash # Check if DeepEP is installed python3 -c "import deep_ep; print('✓ DeepEP installed')" 2>/dev/null || \ python3 -c "import deepep; print('✓ DeepEP installed')" 2>/dev/null || \ echo "✗ DeepEP not installed (required for MoE models)" ``` If DeepEP is not installed and you need to run MoE models, use the `deepep-installer` skill: ```bash /deepep-installer ``` ## Installation Workflow ### Pre-requisites (Ubuntu 24.04) Ubuntu 24.04 doesn't include pip by default. Install it first: ```bash sudo apt-get update sudo apt-get install -y python3-pip ``` ### Step 1: Environment Setup To set up the environment, ensure CUDA is properly configured: ```bash export CUDA_HOME=/usr/local/cuda export PATH=$CUDA_HOME/bin:$PATH # Set HuggingFace cache to LSSD (if available) if [ -d /lssd/huggingface ]; then export HF_HOME=/lssd/huggingface fi ``` ### Step 2: Install PyTorch ```bash pip install torch torchvision torchaudio \ --index-url https://download.pytorch.org/whl/cu129 ``` ### Step 3: Install vLLM ```bash # Install from PyPI (recommended) pip install vllm==0.14.1 \ --extra-index-url https://download.pytorch.org/whl/cu129 ``` ### Step 4: Install NVIDIA Libraries These libraries must be installed with `--force-reinstall --no-deps` to avoid version conflicts: ```bash pip install nvidia-nccl-cu12==2.28.3 --force-reinstall --no-deps pip install nvidia-cudnn-cu12==9.16.0.29 --force-reinstall --no-deps pip install nvidia-cusparselt-cu12 --force-reinstall --no-deps ``` ### Step 5: Install FlashInfer FlashInfer is the recommended attention backend for vLLM: ```bash pip install flashinfer-python==0.5.3 flashinfer-cubin==0.5.3 ``` **Important:** `flashinfer-python` and `flashinfer-cubin` versions MUST match exactly. **⚠️ WARNING: FlashInfer may change PyTorch version!** FlashInfer installation can upgrade PyTorch from 2.9.1 to 2.10.0, which breaks: - vLLM (requires PyTorch 2.9.1) - sgl-kernel (ABI mismatch) - DeepEP (ABI mismatch) **Always reinstall after FlashInfer:** ```bash pip install flashinfer-python==0.5.3 flashinfer-cubin==0.5.3 pip install torch==2.9.1+cu129 --index-url https://download.pytorch.org/whl/cu129 --force-reinstall pip install nvidia-nccl-cu12==2.28.3 nvidia-cudnn-cu12==9.16.0.29 --force-reinstall --no-deps ``` ### Step 6: Configure LD_LIBRARY_PATH To fix library loading issues, run `scripts/setup_env.sh` or manually set: ```bash # Collect all nvidia pip package lib paths NVIDIA_LIB_PATHS="" for d in /usr/local/lib/python3.*/dist-packages/nvidia/*/lib; do [ -d "$d" ] && NVIDIA_LIB_PATHS="${d}:${NVIDIA_LIB_PATHS}" done for d in $HOME/.local/lib/python3.*/site-packages/nvidia/*/lib; do [ -d "$d" ] && NVIDIA_LIB_PATHS="${d}:${NVIDIA_LIB_PATHS}" done export LD_LIBRARY_PATH=${CUDA_HOME}/lib64:${NVIDIA_LIB_PATHS}${LD_LIBRARY_PATH} ``` ## Common Errors and Fixes ### Error: libcudnn.so.9 not found **Symptom:** ``` ImportError: libcudnn.so.9: cannot open shared object file ``` **Fix:** ```bash pip install nvidia-cudnn-cu12==9.16.0.29 --force-reinstall --no-deps # Then set LD_LIBRARY_PATH as described above ``` ### Error: libcusparseLt.so.0 not found **Symptom:** ``` ImportError: libcusparseLt.so.0: cannot open shared object file ``` **Fix:** ```bash pip install nvidia-cusparselt-cu12 --force-reinstall --no-deps # Then set LD_LIBRARY_PATH as described above ``` ### Error: FlashInfer version mismatch **Symptom:** ``` ModuleNotFoundError: No module named 'flashinfer.jit.cubin_loader' ``` or ``` FLASHINFER_CUBIN_DIR not found ``` **Diagnosis:** `flashinfer-python` and `flashinfer-cubin` versions don't match. **Fix:** ```bash pip install flashinfer-python==0.5.3 flashinfer-cubin==0.5.3 --force-reinstall ``` ### Error: WorkerProc failed to start **Symptom:** ``` ERROR: WorkerProc failed to start. File "vllm/v1/attention/selector.py" ... ``` **Diagnosis:** Usually caused by FlashInfer import failure. **Fix:** Check FlashInfer versions match and LD_LIBRARY_PATH is set correctly. ### Error: assert self.total_num_heads % tp_size == 0 **Symptom:** ``` AssertionError: assert self.total_num_heads % tp_size == 0 ``` **Diagnosis:** The model's attention head count is not divisible by the tensor parallelism size. **Fix:** Choose a `--tensor-parallel-size` value that divides the model's attention head count: | Model | Attention Heads | Valid TP Values | |-------|-----------------|-----------------| | Qwen2.5-7B | 28 | 1, 2, 4, 7, 14 | | Qwen2.5-72B | 64 | 1, 2, 4, 8, 16, 32 | | Llama-3-8B | 32 | 1, 2, 4, 8, 16, 32 | | Llama-3-70B | 64 | 1, 2, 4, 8, 16, 32 | | DeepSeek-R1 | 128 | 1, 2, 4, 8, 16, 32, 64 | To find the attention head count for any model: ```bash python3 -c "from transformers import AutoConfig; c = AutoConfig.from_pretrained('MODEL_NAME'); print(f'Attention heads: {c.num_attention_heads}')" ``` ## Starting the Server To start the vLLM OpenAI-compatible API server: ```bash # Load environment source /vllm-workspace/vllm-env.sh # Start server (adjust tp based on model architecture) vllm serve Qwen/Qwen2.5-7B-Instruct \ --tensor-parallel-size 4 \ --port 8000 \ --host 0.0.0.0 ``` Or using the Python module: ```bash python3 -m vllm.entrypoints.openai.api_server \ --model Qwen/Qwen2.5-7B-Instruct \ --tensor-parallel-size 4 \ --port 8000 \ --host 0.0.0.0 ``` ## Disaggregated Prefill (PD Separation) vLLM supports prefill-decode disaggregation where prefill and decode phases run on separate instances. This allows independent tuning of TTFT (time-to-first-token) and ITL (inter-token-latency). ### KV Transfer Connectors vLLM supports multiple KV transfer backends: | Connector | Dependency | Use Case | |-----------|------------|----------| | **NixlConnector** | nixl | Recommended, uses DMA-BUF (no nvidia_peermem needed) | | MooncakeConnector | mooncake | Requires nvidia_peermem kernel module | | P2pNcclConnector | NCCL | Same-node P2P transfer | | LMCacheConnector | lmcache | External KV cache | ### Installing NIXL for Disaggregation ```bash pip install --break-system-packages nixl==0.9.0 # IMPORTANT: NIXL may downgrade NVIDIA libraries, reinstall correct versions: pip install nvidia-nccl-cu12==2.28.3 --force-reinstall --no-deps pip install nvidia-cudnn-cu12==9.16.0.29 --force-reinstall --no-deps ``` ### Verifying NIXL ```bash python3 -c " from vllm.distributed.kv_transfer.kv_connector.v1.nixl_connector import NixlConnector print('NixlConnector OK') " ``` ### Disaggregated Prefill Configuration **Prefill Node:** ```bash vllm serve deepseek-ai/DeepSeek-V3 \ --tensor-parallel-size 8 \ --port 8100 \ --download-dir /lssd/huggingface/hub \ --kv-transfer-config '{ "kv_connector": "NixlConnector", "kv_role": "kv_both", "kv_buffer_device": "cuda" }' ``` **Decode Node:** ```bash vllm serve deepseek-ai/DeepSeek-V3 \ --tensor-parallel-size 8 \ --port 8200 \ --download-dir /lssd/huggingface/hub \ --kv-transfer-config '{ "kv_connector": "NixlConnector", "kv_role": "kv_both", "kv_buffer_device": "cuda" }' ``` ### KV Transfer Config Parameters | Parameter | Values | Description | |-----------|--------|-------------| | kv_connector | NixlConnector, MooncakeConnector, P2pNcclConnector | KV transfer backend | | kv_role | kv_producer, kv_consumer, kv_both | Role in KV transfer (kv_both for most cases) | | kv_buffer_device | cuda, cpu | Buffer device (cuda recommended, cpu for TPU) | | kv_ip | IP address | Connector IP for distributed connection | | kv_port | Port number | Connector port (default: 14579) | ### NIXL vs Mooncake | Feature | NIXL | Mooncake | |---------|------|----------|
Auf GitHub ansehen
Diese SKILL.md ist sehr gross, daher zeigt SkillsMP hier nur den ersten Abschnitt. Auf GitHub ansehen