- name
- vllm-installer
- description
- This skill should be used when users need to install, configure, debug, or run vLLM inference server on NVIDIA GPUs (especially B200/H100/A100). It covers installation from PyPI or source, dependency management, environment setup, common error diagnosis and fixes, tensor parallelism configuration, and server startup/testing. The skill automatically checks for LSSD mount status and DeepEP installation for MoE models.
- license
- MIT
# vLLM Installer
This skill provides comprehensive guidance for installing, configuring, and debugging vLLM on NVIDIA GPUs with CUDA 12.x.
## When to Use This Skill
- Installing vLLM on NVIDIA GPUs (B200/H100/A100)
- Debugging vLLM installation errors (missing libraries, version conflicts)
- Configuring tensor parallelism for different model architectures
- Setting up environment variables for CUDA and NVIDIA libraries
- Starting and testing vLLM OpenAI-compatible API server
- Fixing common runtime errors (cuDNN, cusparseLt, FlashInfer issues)
## Version Information (as of v0.14.1)
| Component | Version | Notes |
|-----------|---------|-------|
| vLLM | 0.14.1 | Latest stable (v0.15.0rc2 not yet on PyPI) |
| flashinfer-python | 0.5.3 | Attention backend |
| flashinfer-cubin | 0.5.3 | Must match flashinfer-python version |
| nixl | 0.9.0 | KV cache transfer (DMA-BUF, recommended for PD disaggregation) |
| nvidia-nccl-cu12 | 2.28.3 | Force reinstall |
| nvidia-cudnn-cu12 | 9.16.0.29 | Required for PyTorch 2.9+ |
| bitsandbytes | 0.46.1 | Quantization support |
| numpy | <2.3 | Required for numba compatibility |
### DeepSeek-V3 FP8 on Blackwell (B200)
**vLLM correctly handles DeepSeek-V3 FP8 on Blackwell GPUs**, unlike SGLang which produces garbage output due to FP8 scale format incompatibility.
| Framework | DeepSeek-V3 FP8 on Blackwell |
|-----------|------------------------------|
| **vLLM 0.14.1** | ✅ Works correctly |
| SGLang 0.5.8 | ❌ Garbage output (scale format mismatch) |
If you need to run DeepSeek-V3 on Blackwell (B200) GPUs, **use vLLM**:
```bash
vllm serve deepseek-ai/DeepSeek-V3 \
--tensor-parallel-size 8 \
--port 8000 \
--download-dir /lssd/huggingface/hub \
--trust-remote-code \
--max-model-len 4096
```
## Pre-Installation Checks
### Step 0: Check Prerequisites
Before installing vLLM, the skill automatically checks:
1. **LSSD Mount Status** - High-speed local SSD for model caching
2. **DeepEP Installation** - Required for MoE models (DeepSeek-V3, DeepSeek-R1)
#### LSSD Check
```bash
# Check if /lssd is mounted
if mountpoint -q /lssd 2>/dev/null; then
echo "✓ LSSD is mounted: $(df -h /lssd | tail -1 | awk '{print $2}')"
else
echo "✗ LSSD is not mounted"
echo " Run: /lssd-mounter"
fi
```
If LSSD is not mounted, use the `lssd-mounter` skill:
```bash
/lssd-mounter
```
#### DeepEP Check (for MoE models)
```bash
# Check if DeepEP is installed
python3 -c "import deep_ep; print('✓ DeepEP installed')" 2>/dev/null || \
python3 -c "import deepep; print('✓ DeepEP installed')" 2>/dev/null || \
echo "✗ DeepEP not installed (required for MoE models)"
```
If DeepEP is not installed and you need to run MoE models, use the `deepep-installer` skill:
```bash
/deepep-installer
```
## Installation Workflow
### Pre-requisites (Ubuntu 24.04)
Ubuntu 24.04 doesn't include pip by default. Install it first:
```bash
sudo apt-get update
sudo apt-get install -y python3-pip
```
### Step 1: Environment Setup
To set up the environment, ensure CUDA is properly configured:
```bash
export CUDA_HOME=/usr/local/cuda
export PATH=$CUDA_HOME/bin:$PATH
# Set HuggingFace cache to LSSD (if available)
if [ -d /lssd/huggingface ]; then
export HF_HOME=/lssd/huggingface
fi
```
### Step 2: Install PyTorch
```bash
pip install torch torchvision torchaudio \
--index-url https://download.pytorch.org/whl/cu129
```
### Step 3: Install vLLM
```bash
# Install from PyPI (recommended)
pip install vllm==0.14.1 \
--extra-index-url https://download.pytorch.org/whl/cu129
```
### Step 4: Install NVIDIA Libraries
These libraries must be installed with `--force-reinstall --no-deps` to avoid version conflicts:
```bash
pip install nvidia-nccl-cu12==2.28.3 --force-reinstall --no-deps
pip install nvidia-cudnn-cu12==9.16.0.29 --force-reinstall --no-deps
pip install nvidia-cusparselt-cu12 --force-reinstall --no-deps
```
### Step 5: Install FlashInfer
FlashInfer is the recommended attention backend for vLLM:
```bash
pip install flashinfer-python==0.5.3 flashinfer-cubin==0.5.3
```
**Important:** `flashinfer-python` and `flashinfer-cubin` versions MUST match exactly.
**⚠️ WARNING: FlashInfer may change PyTorch version!**
FlashInfer installation can upgrade PyTorch from 2.9.1 to 2.10.0, which breaks:
- vLLM (requires PyTorch 2.9.1)
- sgl-kernel (ABI mismatch)
- DeepEP (ABI mismatch)
**Always reinstall after FlashInfer:**
```bash
pip install flashinfer-python==0.5.3 flashinfer-cubin==0.5.3
pip install torch==2.9.1+cu129 --index-url https://download.pytorch.org/whl/cu129 --force-reinstall
pip install nvidia-nccl-cu12==2.28.3 nvidia-cudnn-cu12==9.16.0.29 --force-reinstall --no-deps
```
### Step 6: Configure LD_LIBRARY_PATH
To fix library loading issues, run `scripts/setup_env.sh` or manually set:
```bash
# Collect all nvidia pip package lib paths
NVIDIA_LIB_PATHS=""
for d in /usr/local/lib/python3.*/dist-packages/nvidia/*/lib; do
[ -d "$d" ] && NVIDIA_LIB_PATHS="${d}:${NVIDIA_LIB_PATHS}"
done
for d in $HOME/.local/lib/python3.*/site-packages/nvidia/*/lib; do
[ -d "$d" ] && NVIDIA_LIB_PATHS="${d}:${NVIDIA_LIB_PATHS}"
done
export LD_LIBRARY_PATH=${CUDA_HOME}/lib64:${NVIDIA_LIB_PATHS}${LD_LIBRARY_PATH}
```
## Common Errors and Fixes
### Error: libcudnn.so.9 not found
**Symptom:**
```
ImportError: libcudnn.so.9: cannot open shared object file
```
**Fix:**
```bash
pip install nvidia-cudnn-cu12==9.16.0.29 --force-reinstall --no-deps
# Then set LD_LIBRARY_PATH as described above
```
### Error: libcusparseLt.so.0 not found
**Symptom:**
```
ImportError: libcusparseLt.so.0: cannot open shared object file
```
**Fix:**
```bash
pip install nvidia-cusparselt-cu12 --force-reinstall --no-deps
# Then set LD_LIBRARY_PATH as described above
```
### Error: FlashInfer version mismatch
**Symptom:**
```
ModuleNotFoundError: No module named 'flashinfer.jit.cubin_loader'
```
or
```
FLASHINFER_CUBIN_DIR not found
```
**Diagnosis:** `flashinfer-python` and `flashinfer-cubin` versions don't match.
**Fix:**
```bash
pip install flashinfer-python==0.5.3 flashinfer-cubin==0.5.3 --force-reinstall
```
### Error: WorkerProc failed to start
**Symptom:**
```
ERROR: WorkerProc failed to start.
File "vllm/v1/attention/selector.py" ...
```
**Diagnosis:** Usually caused by FlashInfer import failure.
**Fix:** Check FlashInfer versions match and LD_LIBRARY_PATH is set correctly.
### Error: assert self.total_num_heads % tp_size == 0
**Symptom:**
```
AssertionError: assert self.total_num_heads % tp_size == 0
```
**Diagnosis:** The model's attention head count is not divisible by the tensor parallelism size.
**Fix:** Choose a `--tensor-parallel-size` value that divides the model's attention head count:
| Model | Attention Heads | Valid TP Values |
|-------|-----------------|-----------------|
| Qwen2.5-7B | 28 | 1, 2, 4, 7, 14 |
| Qwen2.5-72B | 64 | 1, 2, 4, 8, 16, 32 |
| Llama-3-8B | 32 | 1, 2, 4, 8, 16, 32 |
| Llama-3-70B | 64 | 1, 2, 4, 8, 16, 32 |
| DeepSeek-R1 | 128 | 1, 2, 4, 8, 16, 32, 64 |
To find the attention head count for any model:
```bash
python3 -c "from transformers import AutoConfig; c = AutoConfig.from_pretrained('MODEL_NAME'); print(f'Attention heads: {c.num_attention_heads}')"
```
## Starting the Server
To start the vLLM OpenAI-compatible API server:
```bash
# Load environment
source /vllm-workspace/vllm-env.sh
# Start server (adjust tp based on model architecture)
vllm serve Qwen/Qwen2.5-7B-Instruct \
--tensor-parallel-size 4 \
--port 8000 \
--host 0.0.0.0
```
Or using the Python module:
```bash
python3 -m vllm.entrypoints.openai.api_server \
--model Qwen/Qwen2.5-7B-Instruct \
--tensor-parallel-size 4 \
--port 8000 \
--host 0.0.0.0
```
## Disaggregated Prefill (PD Separation)
vLLM supports prefill-decode disaggregation where prefill and decode phases run on separate instances. This allows independent tuning of TTFT (time-to-first-token) and ITL (inter-token-latency).
### KV Transfer Connectors
vLLM supports multiple KV transfer backends:
| Connector | Dependency | Use Case |
|-----------|------------|----------|
| **NixlConnector** | nixl | Recommended, uses DMA-BUF (no nvidia_peermem needed) |
| MooncakeConnector | mooncake | Requires nvidia_peermem kernel module |
| P2pNcclConnector | NCCL | Same-node P2P transfer |
| LMCacheConnector | lmcache | External KV cache |
### Installing NIXL for Disaggregation
```bash
pip install --break-system-packages nixl==0.9.0
# IMPORTANT: NIXL may downgrade NVIDIA libraries, reinstall correct versions:
pip install nvidia-nccl-cu12==2.28.3 --force-reinstall --no-deps
pip install nvidia-cudnn-cu12==9.16.0.29 --force-reinstall --no-deps
```
### Verifying NIXL
```bash
python3 -c "
from vllm.distributed.kv_transfer.kv_connector.v1.nixl_connector import NixlConnector
print('NixlConnector OK')
"
```
### Disaggregated Prefill Configuration
**Prefill Node:**
```bash
vllm serve deepseek-ai/DeepSeek-V3 \
--tensor-parallel-size 8 \
--port 8100 \
--download-dir /lssd/huggingface/hub \
--kv-transfer-config '{
"kv_connector": "NixlConnector",
"kv_role": "kv_both",
"kv_buffer_device": "cuda"
}'
```
**Decode Node:**
```bash
vllm serve deepseek-ai/DeepSeek-V3 \
--tensor-parallel-size 8 \
--port 8200 \
--download-dir /lssd/huggingface/hub \
--kv-transfer-config '{
"kv_connector": "NixlConnector",
"kv_role": "kv_both",
"kv_buffer_device": "cuda"
}'
```
### KV Transfer Config Parameters
| Parameter | Values | Description |
|-----------|--------|-------------|
| kv_connector | NixlConnector, MooncakeConnector, P2pNcclConnector | KV transfer backend |
| kv_role | kv_producer, kv_consumer, kv_both | Role in KV transfer (kv_both for most cases) |
| kv_buffer_device | cuda, cpu | Buffer device (cuda recommended, cpu for TPU) |
| kv_ip | IP address | Connector IP for distributed connection |
| kv_port | Port number | Connector port (default: 14579) |
### NIXL vs Mooncake
| Feature | NIXL | Mooncake |
|---------|------|----------|
Auf GitHub ansehen