Skip to main content

sglang-installer

This skill should be used when users need to install, configure, debug, or run SGLang inference server on NVIDIA GPUs (especially B200/H100/A100). It covers installation from source, dependency management, environment setup, common error diagnosis and fixes, tensor parallelism configuration, and server startup/testing.

Jump to install

Source facts

Repository
yangwhale/gpu-tpu-pedia
Last source activity
February 9, 2026 at 09:12
Detected SKILL.md language
Mixed languages
Stars
14
Forks
3

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.

File Explorer
5 files

Showing SKILL.md

SKILL.md
Source instructions · Read-only preview
name
sglang-installer
description
This skill should be used when users need to install, configure, debug, or run SGLang inference server on NVIDIA GPUs (especially B200/H100/A100). It covers installation from source, dependency management, environment setup, common error diagnosis and fixes, tensor parallelism configuration, and server startup/testing.
license
MIT
# SGLang Installer This skill provides comprehensive guidance for installing, configuring, and debugging SGLang on NVIDIA GPUs with CUDA 12.x. ## When to Use This Skill - Installing SGLang from source on NVIDIA GPUs - Debugging SGLang installation errors (missing libraries, version conflicts) - Configuring tensor parallelism for different model architectures - Setting up environment variables for CUDA and NVIDIA libraries - Starting and testing SGLang inference server - Fixing common runtime errors (cuDNN, cusparseLt, NCCL issues) ## Version Information (as of v0.5.8) | Component | Version | Notes | |-----------|---------|-------| | SGLang | 0.5.8 | Latest stable (2026-01-29) | | sgl-kernel | 0.3.21 | PyPI install for CUDA 12.9 | | mooncake-transfer-engine | 0.3.8.post1 | KV cache transfer (requires nvidia_peermem) | | nixl | 0.9.0 | KV cache transfer (DMA-BUF, recommended) | | nvidia-nccl-cu12 | 2.28.3 | Force reinstall | | nvidia-cudnn-cu12 | 9.16.0.29 | Required for PyTorch 2.9+ | | flashinfer | 0.6.1 | SGLang 0.5.8 requires 0.6.1 (vLLM uses 0.5.3) | | sglang-router | 0.5.8 | PD disaggregation 路由 (pip install sglang-router) | ### What's New in v0.5.8 - **1.5x faster diffusion models** across the board - **Chunked Pipeline Parallelism** for million-token context (near-linear scaling) - **EPD Disaggregation** for Vision-Language Models (elastic encoder scaling) - **GLM4-MoE optimization**: 65% faster TTFT - **New models**: GLM 4.7 Flash, LFM2, Qwen3-VL-Embedding/Reranker, DeepSeek V3.2 NVFP4, FLUX.2-klein-9B ### What's New in v0.5.7 - **Model Gateway v0.3.0** release - **Scalable Pipeline Parallelism** with dynamic chunking for ultra-long contexts - **Encoder Disaggregation** for multi-modal models - **Diffusion improvements**: `--dit-layerwise-offload true` reduces peak VRAM by 30GB - **New models**: Mimo-V2-Flash, Nemotron-Nano-v3, LLaDA 2.0, EAGLE 3 speculative decoding - **Hardware support**: AMD/4090/5090 for diffusion ## Mooncake Transfer Engine Mooncake is required for **prefill-decode disaggregation** mode, which separates prefill and decode phases across different nodes for production deployments. ### Installing Mooncake ```bash pip install --break-system-packages mooncake-transfer-engine==0.3.8.post1 ``` ### Verifying Mooncake ```bash python3 -c "from mooncake.engine import TransferEngine; print('Mooncake OK')" ``` ### When is Mooncake Needed? Mooncake is required when using: - `--disaggregation-mode prefill` or `--disaggregation-mode decode` - Multi-node deployments with KV cache transfer - Production DeepSeek-V3/R1 deployments with prefill-decode separation **Note:** For single-node testing without disaggregation, Mooncake is not required. ## NIXL Transfer Engine (Recommended) NIXL (NVIDIA Inference Xfer Library) is an alternative to Mooncake that uses **DMA-BUF** instead of nvidia_peermem. It's the recommended choice when: - Using NVIDIA Open Kernel Module (nvidia_peermem won't load) - nvidia_peermem fails with "Invalid argument" error - You want a more portable solution that doesn't depend on kernel modules ### Installing NIXL ```bash pip install --break-system-packages nixl==0.9.0 # IMPORTANT: NIXL may downgrade NVIDIA libraries, reinstall correct versions: pip install nvidia-nccl-cu12==2.28.3 --force-reinstall --no-deps pip install nvidia-cudnn-cu12==9.16.0.29 --force-reinstall --no-deps ``` **注意**: NIXL 还会安装 `nvidia-nvshmem-cu12==3.4.5`,这个 pip 包**不会被使用**。 DeepEP 使用的是自编译的 NVSHMEM 3.5.19(带 IBGDA 支持),通过 `unified-env.sh` 中的 `LD_PRELOAD` 加载。 ### Verifying NIXL ```bash python3 -c "import nixl; print('NIXL OK')" ``` ### Using NIXL for Disaggregation Add `--disaggregation-transfer-backend nixl` to your launch command: ```bash python3 -m sglang.launch_server \ --model-path deepseek-ai/DeepSeek-V3 \ --disaggregation-mode prefill \ --disaggregation-transfer-backend nixl \ # Use NIXL instead of Mooncake --tp-size 8 \ ... ``` ### NIXL vs Mooncake | Feature | NIXL | Mooncake | |---------|------|----------| | Memory registration | DMA-BUF (kernel native) | nvidia_peermem (kernel module) | | Transport | UCX (TCP/RDMA/SHM) | RDMA or TCP | | Kernel module required | No | nvidia_peermem (may fail) | | Open Kernel Module compatible | Yes | No (fails to load) | | Recommended for | NVIDIA Open driver, B200 | Legacy systems with nvidia_peermem | **Recommendation:** Use NIXL for new deployments, especially on systems with NVIDIA Open Kernel Module. ## Installation Workflow ### Pre-requisites (Ubuntu 24.04) Ubuntu 24.04 doesn't include pip by default. Install it first: ```bash sudo apt-get update sudo apt-get install -y python3-pip ``` ### Step 1: Environment Setup To set up the environment, ensure CUDA is properly configured: ```bash export CUDA_HOME=/usr/local/cuda export PATH=$CUDA_HOME/bin:$PATH export BUILD_TYPE=blackwell # or "all" for general, "hopper" for H100 ``` ### Step 2: Clone and Install To install SGLang from source: ```bash mkdir -p /sgl-workspace && cd /sgl-workspace # Clone specific version git clone -b v0.5.8 --depth 1 https://github.com/sgl-project/sglang.git cd sglang # Install sgl-kernel first (for CUDA 12.9) pip install sgl-kernel==0.3.21 # Install SGLang with blackwell support pip install -e "python[blackwell]" --extra-index-url https://download.pytorch.org/whl/cu129 ``` ### Step 3: Install Additional Dependencies To install required NVIDIA libraries and NIXL: ```bash # NVIDIA libraries (required) pip install nvidia-nccl-cu12==2.28.3 --force-reinstall --no-deps pip install nvidia-cudnn-cu12==9.16.0.29 --force-reinstall --no-deps # NIXL for KV cache transfer (RECOMMENDED for disaggregation mode) pip install --break-system-packages nixl==0.9.0 # Re-install NVIDIA libs after NIXL (NIXL may downgrade them) pip install nvidia-nccl-cu12==2.28.3 --force-reinstall --no-deps pip install nvidia-cudnn-cu12==9.16.0.29 --force-reinstall --no-deps ``` > ⚠️ **Important**: NIXL is required for prefill-decode disaggregation mode. If you skip NIXL, you'll need nvidia_peermem kernel module (often fails on NVIDIA Open driver). ### Step 4: Configure LD_LIBRARY_PATH To fix library loading issues, run `scripts/setup_env.sh` or manually set: ```bash # Collect all nvidia pip package lib paths NVIDIA_LIB_PATHS="" for d in /usr/local/lib/python3.12/dist-packages/nvidia/*/lib; do [ -d "$d" ] && NVIDIA_LIB_PATHS="${d}:${NVIDIA_LIB_PATHS}" done for d in $HOME/.local/lib/python3.12/site-packages/nvidia/*/lib; do [ -d "$d" ] && NVIDIA_LIB_PATHS="${d}:${NVIDIA_LIB_PATHS}" done export LD_LIBRARY_PATH=${CUDA_HOME}/lib64:${NVIDIA_LIB_PATHS}${LD_LIBRARY_PATH} ``` ## Common Errors and Fixes ### Error: DeepSeek-V3 FP8 outputs garbage on Blackwell (B200) **Symptom:** ``` Model outputs garbage characters like "################" or random symbols ``` **Log Warning:** ``` WARNING model_config.py:872: DeepGemm is enabled but the scale_fmt of checkpoint is not ue8m0. This might cause accuracy degradation on Blackwell. ``` **Diagnosis:** The DeepSeek-V3 official FP8 checkpoint uses `mscale` format which is incompatible with DeepGEMM on Blackwell (B200) GPUs. DeepGEMM expects `ue8m0` scale format. **Workaround:** Use **vLLM** instead of SGLang for DeepSeek-V3 on Blackwell: ```bash # vLLM handles the FP8 scale format correctly vllm serve deepseek-ai/DeepSeek-V3 --tensor-parallel-size 8 --port 8000 --trust-remote-code ``` **Status:** Known issue as of SGLang 0.5.8 on Blackwell GPUs. Works correctly on H100/A100. | Framework | DeepSeek-V3 FP8 on Blackwell | |-----------|------------------------------| | SGLang 0.5.8 | ❌ Garbage output | | vLLM 0.14.1 | ✅ Works correctly | --- ### Error: Cannot uninstall typing_extensions (Ubuntu 24.04) **Symptom:** ``` ERROR: Cannot uninstall typing_extensions 4.10.0, RECORD file not found. Hint: The package was installed by debian. ``` **Diagnosis:** Ubuntu 24.04 installs `typing_extensions` as a system package managed by apt. **Fix:** ```bash pip install -e "python[blackwell]" --ignore-installed typing_extensions --break-system-packages ``` --- ### Error: sgl-kernel ABI mismatch (PyTorch user/system conflict) **Symptom:** ``` ImportError: .../sgl_kernel/sm100/common_ops.abi3.so: undefined symbol: _ZN3c104cuda29c10_cuda_check_implementationEiPKcS2_ib ``` **Diagnosis:** PyTorch installed in both user (`~/.local/lib/python3.12/site-packages/`) and system (`/usr/local/lib/python3.12/dist-packages/`) directories with different versions. **Check:** ```bash pip3 show torch | grep -E "Version|Location" ls /usr/local/lib/python3.12/dist-packages/ | grep torch ls ~/.local/lib/python3.12/site-packages/ | grep torch ``` **Fix:** Remove user-installed torch to use system version: ```bash pip3 uninstall torch torchvision torchaudio -y --break-system-packages python3 -c "import torch; print(torch.__version__)" # Should show system version ``` --- ### Error: libcudnn.so.9 not found **Symptom:** ``` ImportError: libcudnn.so.9: cannot open shared object file ``` **Fix:** ```bash pip install nvidia-cudnn-cu12==9.16.0.29 --force-reinstall --no-deps # Then set LD_LIBRARY_PATH as described above ``` ### Error: libcusparseLt.so.0 not found **Symptom:** ``` ImportError: libcusparseLt.so.0: cannot open shared object file ``` **Fix:** ```bash pip install nvidia-cusparselt-cu12 # Then set LD_LIBRARY_PATH as described above ``` ### Error: assert self.total_num_heads % tp_size == 0 **Symptom:** ``` AssertionError: assert self.total_num_heads % tp_size == 0 ``` **Diagnosis:** The model's attention head count is not divisible by the tensor parallelism size. **Fix:** Choose a `--tp` value that divides the model's attention head count: | Model | Attention Heads | Valid TP Values | |-------|-----------------|-----------------| | Qwen2.5-7B | 28 | 1, 2, 4, 7, 14 | | Qwen2.5-72B | 64 | 1, 2, 4, 8, 16, 32 | | Llama-3-8B | 32 | 1, 2, 4, 8, 16, 32 | | Llama-3-70B | 64 | 1, 2, 4, 8, 16, 32 | | DeepSeek-R1 | 128 | 1, 2, 4, 8, 16, 32, 64 | To find the attention head count for any model: ```bash python3 -c "from transformers import AutoConfig; c = AutoConfig.from_pretrained('MODEL_NAME'); print(f'Attention heads: {c.num_attention_heads}')" ``` ### Error: NCCL errors or timeouts **Fix:** ```bash pip install nvidia-nccl-cu12==2.28.3 --force-reinstall --no-deps ``` ### Error: sgl-kernel version mismatch **Symptom:** SGLang installs an older sgl-kernel version than expected. **Note:** As of v0.5.6.post2, SGLang's dependencies pin sgl-kernel to 0.3.19, so even if you pre-install 0.3.21, it will be downgraded during SGLang installation. This is expected behavior and 0.3.19 works correctly. **If you need a specific version:** Install sgl-kernel AFTER SGLang: ```bash pip install -e "python[blackwell]" ... pip install sgl-kernel==0.3.21 --force-reinstall --no-deps # if needed ``` ### Error: sgl-kernel ABI incompatibility (undefined symbol) **Symptom:** ``` ImportError: .../sgl_kernel/sm100/common_ops.abi3.so: undefined symbol: _ZN3c104cuda29c10_cuda_check_implementationEiPKcS2_ib
View on GitHub
This SKILL.md is very large, so SkillsMP previews the first section here. View on GitHub