- name
- sglang-installer
- description
- This skill should be used when users need to install, configure, debug, or run SGLang inference server on NVIDIA GPUs (especially B200/H100/A100). It covers installation from source, dependency management, environment setup, common error diagnosis and fixes, tensor parallelism configuration, and server startup/testing.
- license
- MIT
# SGLang Installer
This skill provides comprehensive guidance for installing, configuring, and debugging SGLang on NVIDIA GPUs with CUDA 12.x.
## When to Use This Skill
- Installing SGLang from source on NVIDIA GPUs
- Debugging SGLang installation errors (missing libraries, version conflicts)
- Configuring tensor parallelism for different model architectures
- Setting up environment variables for CUDA and NVIDIA libraries
- Starting and testing SGLang inference server
- Fixing common runtime errors (cuDNN, cusparseLt, NCCL issues)
## Version Information (as of v0.5.8)
| Component | Version | Notes |
|-----------|---------|-------|
| SGLang | 0.5.8 | Latest stable (2026-01-29) |
| sgl-kernel | 0.3.21 | PyPI install for CUDA 12.9 |
| mooncake-transfer-engine | 0.3.8.post1 | KV cache transfer (requires nvidia_peermem) |
| nixl | 0.9.0 | KV cache transfer (DMA-BUF, recommended) |
| nvidia-nccl-cu12 | 2.28.3 | Force reinstall |
| nvidia-cudnn-cu12 | 9.16.0.29 | Required for PyTorch 2.9+ |
| flashinfer | 0.6.1 | SGLang 0.5.8 requires 0.6.1 (vLLM uses 0.5.3) |
| sglang-router | 0.5.8 | PD disaggregation 路由 (pip install sglang-router) |
### What's New in v0.5.8
- **1.5x faster diffusion models** across the board
- **Chunked Pipeline Parallelism** for million-token context (near-linear scaling)
- **EPD Disaggregation** for Vision-Language Models (elastic encoder scaling)
- **GLM4-MoE optimization**: 65% faster TTFT
- **New models**: GLM 4.7 Flash, LFM2, Qwen3-VL-Embedding/Reranker, DeepSeek V3.2 NVFP4, FLUX.2-klein-9B
### What's New in v0.5.7
- **Model Gateway v0.3.0** release
- **Scalable Pipeline Parallelism** with dynamic chunking for ultra-long contexts
- **Encoder Disaggregation** for multi-modal models
- **Diffusion improvements**: `--dit-layerwise-offload true` reduces peak VRAM by 30GB
- **New models**: Mimo-V2-Flash, Nemotron-Nano-v3, LLaDA 2.0, EAGLE 3 speculative decoding
- **Hardware support**: AMD/4090/5090 for diffusion
## Mooncake Transfer Engine
Mooncake is required for **prefill-decode disaggregation** mode, which separates prefill and decode phases across different nodes for production deployments.
### Installing Mooncake
```bash
pip install --break-system-packages mooncake-transfer-engine==0.3.8.post1
```
### Verifying Mooncake
```bash
python3 -c "from mooncake.engine import TransferEngine; print('Mooncake OK')"
```
### When is Mooncake Needed?
Mooncake is required when using:
- `--disaggregation-mode prefill` or `--disaggregation-mode decode`
- Multi-node deployments with KV cache transfer
- Production DeepSeek-V3/R1 deployments with prefill-decode separation
**Note:** For single-node testing without disaggregation, Mooncake is not required.
## NIXL Transfer Engine (Recommended)
NIXL (NVIDIA Inference Xfer Library) is an alternative to Mooncake that uses **DMA-BUF** instead of nvidia_peermem. It's the recommended choice when:
- Using NVIDIA Open Kernel Module (nvidia_peermem won't load)
- nvidia_peermem fails with "Invalid argument" error
- You want a more portable solution that doesn't depend on kernel modules
### Installing NIXL
```bash
pip install --break-system-packages nixl==0.9.0
# IMPORTANT: NIXL may downgrade NVIDIA libraries, reinstall correct versions:
pip install nvidia-nccl-cu12==2.28.3 --force-reinstall --no-deps
pip install nvidia-cudnn-cu12==9.16.0.29 --force-reinstall --no-deps
```
**注意**: NIXL 还会安装 `nvidia-nvshmem-cu12==3.4.5`,这个 pip 包**不会被使用**。
DeepEP 使用的是自编译的 NVSHMEM 3.5.19(带 IBGDA 支持),通过 `unified-env.sh` 中的 `LD_PRELOAD` 加载。
### Verifying NIXL
```bash
python3 -c "import nixl; print('NIXL OK')"
```
### Using NIXL for Disaggregation
Add `--disaggregation-transfer-backend nixl` to your launch command:
```bash
python3 -m sglang.launch_server \
--model-path deepseek-ai/DeepSeek-V3 \
--disaggregation-mode prefill \
--disaggregation-transfer-backend nixl \ # Use NIXL instead of Mooncake
--tp-size 8 \
...
```
### NIXL vs Mooncake
| Feature | NIXL | Mooncake |
|---------|------|----------|
| Memory registration | DMA-BUF (kernel native) | nvidia_peermem (kernel module) |
| Transport | UCX (TCP/RDMA/SHM) | RDMA or TCP |
| Kernel module required | No | nvidia_peermem (may fail) |
| Open Kernel Module compatible | Yes | No (fails to load) |
| Recommended for | NVIDIA Open driver, B200 | Legacy systems with nvidia_peermem |
**Recommendation:** Use NIXL for new deployments, especially on systems with NVIDIA Open Kernel Module.
## Installation Workflow
### Pre-requisites (Ubuntu 24.04)
Ubuntu 24.04 doesn't include pip by default. Install it first:
```bash
sudo apt-get update
sudo apt-get install -y python3-pip
```
### Step 1: Environment Setup
To set up the environment, ensure CUDA is properly configured:
```bash
export CUDA_HOME=/usr/local/cuda
export PATH=$CUDA_HOME/bin:$PATH
export BUILD_TYPE=blackwell # or "all" for general, "hopper" for H100
```
### Step 2: Clone and Install
To install SGLang from source:
```bash
mkdir -p /sgl-workspace && cd /sgl-workspace
# Clone specific version
git clone -b v0.5.8 --depth 1 https://github.com/sgl-project/sglang.git
cd sglang
# Install sgl-kernel first (for CUDA 12.9)
pip install sgl-kernel==0.3.21
# Install SGLang with blackwell support
pip install -e "python[blackwell]" --extra-index-url https://download.pytorch.org/whl/cu129
```
### Step 3: Install Additional Dependencies
To install required NVIDIA libraries and NIXL:
```bash
# NVIDIA libraries (required)
pip install nvidia-nccl-cu12==2.28.3 --force-reinstall --no-deps
pip install nvidia-cudnn-cu12==9.16.0.29 --force-reinstall --no-deps
# NIXL for KV cache transfer (RECOMMENDED for disaggregation mode)
pip install --break-system-packages nixl==0.9.0
# Re-install NVIDIA libs after NIXL (NIXL may downgrade them)
pip install nvidia-nccl-cu12==2.28.3 --force-reinstall --no-deps
pip install nvidia-cudnn-cu12==9.16.0.29 --force-reinstall --no-deps
```
> ⚠️ **Important**: NIXL is required for prefill-decode disaggregation mode. If you skip NIXL, you'll need nvidia_peermem kernel module (often fails on NVIDIA Open driver).
### Step 4: Configure LD_LIBRARY_PATH
To fix library loading issues, run `scripts/setup_env.sh` or manually set:
```bash
# Collect all nvidia pip package lib paths
NVIDIA_LIB_PATHS=""
for d in /usr/local/lib/python3.12/dist-packages/nvidia/*/lib; do
[ -d "$d" ] && NVIDIA_LIB_PATHS="${d}:${NVIDIA_LIB_PATHS}"
done
for d in $HOME/.local/lib/python3.12/site-packages/nvidia/*/lib; do
[ -d "$d" ] && NVIDIA_LIB_PATHS="${d}:${NVIDIA_LIB_PATHS}"
done
export LD_LIBRARY_PATH=${CUDA_HOME}/lib64:${NVIDIA_LIB_PATHS}${LD_LIBRARY_PATH}
```
## Common Errors and Fixes
### Error: DeepSeek-V3 FP8 outputs garbage on Blackwell (B200)
**Symptom:**
```
Model outputs garbage characters like "################" or random symbols
```
**Log Warning:**
```
WARNING model_config.py:872: DeepGemm is enabled but the scale_fmt of checkpoint is not ue8m0. This might cause accuracy degradation on Blackwell.
```
**Diagnosis:** The DeepSeek-V3 official FP8 checkpoint uses `mscale` format which is incompatible with DeepGEMM on Blackwell (B200) GPUs. DeepGEMM expects `ue8m0` scale format.
**Workaround:** Use **vLLM** instead of SGLang for DeepSeek-V3 on Blackwell:
```bash
# vLLM handles the FP8 scale format correctly
vllm serve deepseek-ai/DeepSeek-V3 --tensor-parallel-size 8 --port 8000 --trust-remote-code
```
**Status:** Known issue as of SGLang 0.5.8 on Blackwell GPUs. Works correctly on H100/A100.
| Framework | DeepSeek-V3 FP8 on Blackwell |
|-----------|------------------------------|
| SGLang 0.5.8 | ❌ Garbage output |
| vLLM 0.14.1 | ✅ Works correctly |
---
### Error: Cannot uninstall typing_extensions (Ubuntu 24.04)
**Symptom:**
```
ERROR: Cannot uninstall typing_extensions 4.10.0, RECORD file not found.
Hint: The package was installed by debian.
```
**Diagnosis:** Ubuntu 24.04 installs `typing_extensions` as a system package managed by apt.
**Fix:**
```bash
pip install -e "python[blackwell]" --ignore-installed typing_extensions --break-system-packages
```
---
### Error: sgl-kernel ABI mismatch (PyTorch user/system conflict)
**Symptom:**
```
ImportError: .../sgl_kernel/sm100/common_ops.abi3.so: undefined symbol: _ZN3c104cuda29c10_cuda_check_implementationEiPKcS2_ib
```
**Diagnosis:** PyTorch installed in both user (`~/.local/lib/python3.12/site-packages/`) and system (`/usr/local/lib/python3.12/dist-packages/`) directories with different versions.
**Check:**
```bash
pip3 show torch | grep -E "Version|Location"
ls /usr/local/lib/python3.12/dist-packages/ | grep torch
ls ~/.local/lib/python3.12/site-packages/ | grep torch
```
**Fix:** Remove user-installed torch to use system version:
```bash
pip3 uninstall torch torchvision torchaudio -y --break-system-packages
python3 -c "import torch; print(torch.__version__)" # Should show system version
```
---
### Error: libcudnn.so.9 not found
**Symptom:**
```
ImportError: libcudnn.so.9: cannot open shared object file
```
**Fix:**
```bash
pip install nvidia-cudnn-cu12==9.16.0.29 --force-reinstall --no-deps
# Then set LD_LIBRARY_PATH as described above
```
### Error: libcusparseLt.so.0 not found
**Symptom:**
```
ImportError: libcusparseLt.so.0: cannot open shared object file
```
**Fix:**
```bash
pip install nvidia-cusparselt-cu12
# Then set LD_LIBRARY_PATH as described above
```
### Error: assert self.total_num_heads % tp_size == 0
**Symptom:**
```
AssertionError: assert self.total_num_heads % tp_size == 0
```
**Diagnosis:** The model's attention head count is not divisible by the tensor parallelism size.
**Fix:** Choose a `--tp` value that divides the model's attention head count:
| Model | Attention Heads | Valid TP Values |
|-------|-----------------|-----------------|
| Qwen2.5-7B | 28 | 1, 2, 4, 7, 14 |
| Qwen2.5-72B | 64 | 1, 2, 4, 8, 16, 32 |
| Llama-3-8B | 32 | 1, 2, 4, 8, 16, 32 |
| Llama-3-70B | 64 | 1, 2, 4, 8, 16, 32 |
| DeepSeek-R1 | 128 | 1, 2, 4, 8, 16, 32, 64 |
To find the attention head count for any model:
```bash
python3 -c "from transformers import AutoConfig; c = AutoConfig.from_pretrained('MODEL_NAME'); print(f'Attention heads: {c.num_attention_heads}')"
```
### Error: NCCL errors or timeouts
**Fix:**
```bash
pip install nvidia-nccl-cu12==2.28.3 --force-reinstall --no-deps
```
### Error: sgl-kernel version mismatch
**Symptom:** SGLang installs an older sgl-kernel version than expected.
**Note:** As of v0.5.6.post2, SGLang's dependencies pin sgl-kernel to 0.3.19, so even if you pre-install 0.3.21, it will be downgraded during SGLang installation. This is expected behavior and 0.3.19 works correctly.
**If you need a specific version:** Install sgl-kernel AFTER SGLang:
```bash
pip install -e "python[blackwell]" ...
pip install sgl-kernel==0.3.21 --force-reinstall --no-deps # if needed
```
### Error: sgl-kernel ABI incompatibility (undefined symbol)
**Symptom:**
```
ImportError: .../sgl_kernel/sm100/common_ops.abi3.so: undefined symbol: _ZN3c104cuda29c10_cuda_check_implementationEiPKcS2_ib
View on GitHub