Skip to main content

vllm-installer

This skill should be used when users need to install, configure, debug, or run vLLM inference server on NVIDIA GPUs (especially B200/H100/A100). It covers installation from PyPI or source, dependency management, environment setup, common error diagnosis and fixes, tensor parallelism configuration, and server startup/testing. The skill automatically checks for LSSD mount status and DeepEP installation for MoE models.

ソース情報

リポジトリ
yangwhale/gpu-tpu-pedia
ソースの最終更新活動
2026年1月29日 10:42
検出された SKILL.md の言語
英語
スター
14
フォーク
3

インストール方法

デフォルトでは、最初にソースを確認する Prompt が選択されています。直接コマンドに切り替えるか、ローカルコピーをダウンロードすることもできます。

ソースファイルを確認

インストールを決める前に、SKILL.md と SkillsMP に表示されている付属ファイルをお読みください。

ファイルエクスプローラー
5 ファイル

SKILL.md を表示中

SKILL.md
ソースの指示 · 読み取り専用プレビュー
name
vllm-installer
description
This skill should be used when users need to install, configure, debug, or run vLLM inference server on NVIDIA GPUs (especially B200/H100/A100). It covers installation from PyPI or source, dependency management, environment setup, common error diagnosis and fixes, tensor parallelism configuration, and server startup/testing. The skill automatically checks for LSSD mount status and DeepEP installation for MoE models.
license
MIT
# vLLM Installer This skill provides comprehensive guidance for installing, configuring, and debugging vLLM on NVIDIA GPUs with CUDA 12.x. ## When to Use This Skill - Installing vLLM on NVIDIA GPUs (B200/H100/A100) - Debugging vLLM installation errors (missing libraries, version conflicts) - Configuring tensor parallelism for different model architectures - Setting up environment variables for CUDA and NVIDIA libraries - Starting and testing vLLM OpenAI-compatible API server - Fixing common runtime errors (cuDNN, cusparseLt, FlashInfer issues) ## Version Information (as of v0.14.1) | Component | Version | Notes | |-----------|---------|-------| | vLLM | 0.14.1 | Latest stable (v0.15.0rc2 not yet on PyPI) | | flashinfer-python | 0.5.3 | Attention backend | | flashinfer-cubin | 0.5.3 | Must match flashinfer-python version | | nixl | 0.9.0 | KV cache transfer (DMA-BUF, recommended for PD disaggregation) | | nvidia-nccl-cu12 | 2.28.3 | Force reinstall | | nvidia-cudnn-cu12 | 9.16.0.29 | Required for PyTorch 2.9+ | | bitsandbytes | 0.46.1 | Quantization support | | numpy | <2.3 | Required for numba compatibility | ### DeepSeek-V3 FP8 on Blackwell (B200) **vLLM correctly handles DeepSeek-V3 FP8 on Blackwell GPUs**, unlike SGLang which produces garbage output due to FP8 scale format incompatibility. | Framework | DeepSeek-V3 FP8 on Blackwell | |-----------|------------------------------| | **vLLM 0.14.1** | ✅ Works correctly | | SGLang 0.5.8 | ❌ Garbage output (scale format mismatch) | If you need to run DeepSeek-V3 on Blackwell (B200) GPUs, **use vLLM**: ```bash vllm serve deepseek-ai/DeepSeek-V3 \ --tensor-parallel-size 8 \ --port 8000 \ --download-dir /lssd/huggingface/hub \ --trust-remote-code \ --max-model-len 4096 ``` ## Pre-Installation Checks ### Step 0: Check Prerequisites Before installing vLLM, the skill automatically checks: 1. **LSSD Mount Status** - High-speed local SSD for model caching 2. **DeepEP Installation** - Required for MoE models (DeepSeek-V3, DeepSeek-R1) #### LSSD Check ```bash # Check if /lssd is mounted if mountpoint -q /lssd 2>/dev/null; then echo "✓ LSSD is mounted: $(df -h /lssd | tail -1 | awk '{print $2}')" else echo "✗ LSSD is not mounted" echo " Run: /lssd-mounter" fi ``` If LSSD is not mounted, use the `lssd-mounter` skill: ```bash /lssd-mounter ``` #### DeepEP Check (for MoE models) ```bash # Check if DeepEP is installed python3 -c "import deep_ep; print('✓ DeepEP installed')" 2>/dev/null || \ python3 -c "import deepep; print('✓ DeepEP installed')" 2>/dev/null || \ echo "✗ DeepEP not installed (required for MoE models)" ``` If DeepEP is not installed and you need to run MoE models, use the `deepep-installer` skill: ```bash /deepep-installer ``` ## Installation Workflow ### Pre-requisites (Ubuntu 24.04) Ubuntu 24.04 doesn't include pip by default. Install it first: ```bash sudo apt-get update sudo apt-get install -y python3-pip ``` ### Step 1: Environment Setup To set up the environment, ensure CUDA is properly configured: ```bash export CUDA_HOME=/usr/local/cuda export PATH=$CUDA_HOME/bin:$PATH # Set HuggingFace cache to LSSD (if available) if [ -d /lssd/huggingface ]; then export HF_HOME=/lssd/huggingface fi ``` ### Step 2: Install PyTorch ```bash pip install torch torchvision torchaudio \ --index-url https://download.pytorch.org/whl/cu129 ``` ### Step 3: Install vLLM ```bash # Install from PyPI (recommended) pip install vllm==0.14.1 \ --extra-index-url https://download.pytorch.org/whl/cu129 ``` ### Step 4: Install NVIDIA Libraries These libraries must be installed with `--force-reinstall --no-deps` to avoid version conflicts: ```bash pip install nvidia-nccl-cu12==2.28.3 --force-reinstall --no-deps pip install nvidia-cudnn-cu12==9.16.0.29 --force-reinstall --no-deps pip install nvidia-cusparselt-cu12 --force-reinstall --no-deps ``` ### Step 5: Install FlashInfer FlashInfer is the recommended attention backend for vLLM: ```bash pip install flashinfer-python==0.5.3 flashinfer-cubin==0.5.3 ``` **Important:** `flashinfer-python` and `flashinfer-cubin` versions MUST match exactly. **⚠️ WARNING: FlashInfer may change PyTorch version!** FlashInfer installation can upgrade PyTorch from 2.9.1 to 2.10.0, which breaks: - vLLM (requires PyTorch 2.9.1) - sgl-kernel (ABI mismatch) - DeepEP (ABI mismatch) **Always reinstall after FlashInfer:** ```bash pip install flashinfer-python==0.5.3 flashinfer-cubin==0.5.3 pip install torch==2.9.1+cu129 --index-url https://download.pytorch.org/whl/cu129 --force-reinstall pip install nvidia-nccl-cu12==2.28.3 nvidia-cudnn-cu12==9.16.0.29 --force-reinstall --no-deps ``` ### Step 6: Configure LD_LIBRARY_PATH To fix library loading issues, run `scripts/setup_env.sh` or manually set: ```bash # Collect all nvidia pip package lib paths NVIDIA_LIB_PATHS="" for d in /usr/local/lib/python3.*/dist-packages/nvidia/*/lib; do [ -d "$d" ] && NVIDIA_LIB_PATHS="${d}:${NVIDIA_LIB_PATHS}" done for d in $HOME/.local/lib/python3.*/site-packages/nvidia/*/lib; do [ -d "$d" ] && NVIDIA_LIB_PATHS="${d}:${NVIDIA_LIB_PATHS}" done export LD_LIBRARY_PATH=${CUDA_HOME}/lib64:${NVIDIA_LIB_PATHS}${LD_LIBRARY_PATH} ``` ## Common Errors and Fixes ### Error: libcudnn.so.9 not found **Symptom:** ``` ImportError: libcudnn.so.9: cannot open shared object file ``` **Fix:** ```bash pip install nvidia-cudnn-cu12==9.16.0.29 --force-reinstall --no-deps # Then set LD_LIBRARY_PATH as described above ``` ### Error: libcusparseLt.so.0 not found **Symptom:** ``` ImportError: libcusparseLt.so.0: cannot open shared object file ``` **Fix:** ```bash pip install nvidia-cusparselt-cu12 --force-reinstall --no-deps # Then set LD_LIBRARY_PATH as described above ``` ### Error: FlashInfer version mismatch **Symptom:** ``` ModuleNotFoundError: No module named 'flashinfer.jit.cubin_loader' ``` or ``` FLASHINFER_CUBIN_DIR not found ``` **Diagnosis:** `flashinfer-python` and `flashinfer-cubin` versions don't match. **Fix:** ```bash pip install flashinfer-python==0.5.3 flashinfer-cubin==0.5.3 --force-reinstall ``` ### Error: WorkerProc failed to start **Symptom:** ``` ERROR: WorkerProc failed to start. File "vllm/v1/attention/selector.py" ... ``` **Diagnosis:** Usually caused by FlashInfer import failure. **Fix:** Check FlashInfer versions match and LD_LIBRARY_PATH is set correctly. ### Error: assert self.total_num_heads % tp_size == 0 **Symptom:** ``` AssertionError: assert self.total_num_heads % tp_size == 0 ``` **Diagnosis:** The model's attention head count is not divisible by the tensor parallelism size. **Fix:** Choose a `--tensor-parallel-size` value that divides the model's attention head count: | Model | Attention Heads | Valid TP Values | |-------|-----------------|-----------------| | Qwen2.5-7B | 28 | 1, 2, 4, 7, 14 | | Qwen2.5-72B | 64 | 1, 2, 4, 8, 16, 32 | | Llama-3-8B | 32 | 1, 2, 4, 8, 16, 32 | | Llama-3-70B | 64 | 1, 2, 4, 8, 16, 32 | | DeepSeek-R1 | 128 | 1, 2, 4, 8, 16, 32, 64 | To find the attention head count for any model: ```bash python3 -c "from transformers import AutoConfig; c = AutoConfig.from_pretrained('MODEL_NAME'); print(f'Attention heads: {c.num_attention_heads}')" ``` ## Starting the Server To start the vLLM OpenAI-compatible API server: ```bash # Load environment source /vllm-workspace/vllm-env.sh # Start server (adjust tp based on model architecture) vllm serve Qwen/Qwen2.5-7B-Instruct \ --tensor-parallel-size 4 \ --port 8000 \ --host 0.0.0.0 ``` Or using the Python module: ```bash python3 -m vllm.entrypoints.openai.api_server \ --model Qwen/Qwen2.5-7B-Instruct \ --tensor-parallel-size 4 \ --port 8000 \ --host 0.0.0.0 ``` ## Disaggregated Prefill (PD Separation) vLLM supports prefill-decode disaggregation where prefill and decode phases run on separate instances. This allows independent tuning of TTFT (time-to-first-token) and ITL (inter-token-latency). ### KV Transfer Connectors vLLM supports multiple KV transfer backends: | Connector | Dependency | Use Case | |-----------|------------|----------| | **NixlConnector** | nixl | Recommended, uses DMA-BUF (no nvidia_peermem needed) | | MooncakeConnector | mooncake | Requires nvidia_peermem kernel module | | P2pNcclConnector | NCCL | Same-node P2P transfer | | LMCacheConnector | lmcache | External KV cache | ### Installing NIXL for Disaggregation ```bash pip install --break-system-packages nixl==0.9.0 # IMPORTANT: NIXL may downgrade NVIDIA libraries, reinstall correct versions: pip install nvidia-nccl-cu12==2.28.3 --force-reinstall --no-deps pip install nvidia-cudnn-cu12==9.16.0.29 --force-reinstall --no-deps ``` ### Verifying NIXL ```bash python3 -c " from vllm.distributed.kv_transfer.kv_connector.v1.nixl_connector import NixlConnector print('NixlConnector OK') " ``` ### Disaggregated Prefill Configuration **Prefill Node:** ```bash vllm serve deepseek-ai/DeepSeek-V3 \ --tensor-parallel-size 8 \ --port 8100 \ --download-dir /lssd/huggingface/hub \ --kv-transfer-config '{ "kv_connector": "NixlConnector", "kv_role": "kv_both", "kv_buffer_device": "cuda" }' ``` **Decode Node:** ```bash vllm serve deepseek-ai/DeepSeek-V3 \ --tensor-parallel-size 8 \ --port 8200 \ --download-dir /lssd/huggingface/hub \ --kv-transfer-config '{ "kv_connector": "NixlConnector", "kv_role": "kv_both", "kv_buffer_device": "cuda" }' ``` ### KV Transfer Config Parameters | Parameter | Values | Description | |-----------|--------|-------------| | kv_connector | NixlConnector, MooncakeConnector, P2pNcclConnector | KV transfer backend | | kv_role | kv_producer, kv_consumer, kv_both | Role in KV transfer (kv_both for most cases) | | kv_buffer_device | cuda, cpu | Buffer device (cuda recommended, cpu for TPU) | | kv_ip | IP address | Connector IP for distributed connection | | kv_port | Port number | Connector port (default: 14579) | ### NIXL vs Mooncake | Feature | NIXL | Mooncake | |---------|------|----------|
GitHubで見る
この SKILL.md は非常に大きいため、SkillsMP では最初のセクションだけを表示しています。 GitHubで見る