Skip to main content

sglang-installer

This skill should be used when users need to install, configure, debug, or run SGLang inference server on NVIDIA GPUs (especially B200/H100/A100). It covers installation from source, dependency management, environment setup, common error diagnosis and fixes, tensor parallelism configuration, and server startup/testing.

معلومات المصدر

المستودع
yangwhale/gpu-tpu-pedia
آخر نشاط في المصدر
٩ فبراير ٢٠٢٦ في ٠٩:١٢
لغة SKILL.md المكتشفة
لغات متعددة
النجوم
١٤
التفرعات
٣

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.

مستكشف الملفات
5 ملفات

عرض SKILL.md

SKILL.md
تعليمات المصدر · معاينة للقراءة فقط
name
sglang-installer
description
This skill should be used when users need to install, configure, debug, or run SGLang inference server on NVIDIA GPUs (especially B200/H100/A100). It covers installation from source, dependency management, environment setup, common error diagnosis and fixes, tensor parallelism configuration, and server startup/testing.
license
MIT
# SGLang Installer This skill provides comprehensive guidance for installing, configuring, and debugging SGLang on NVIDIA GPUs with CUDA 12.x. ## When to Use This Skill - Installing SGLang from source on NVIDIA GPUs - Debugging SGLang installation errors (missing libraries, version conflicts) - Configuring tensor parallelism for different model architectures - Setting up environment variables for CUDA and NVIDIA libraries - Starting and testing SGLang inference server - Fixing common runtime errors (cuDNN, cusparseLt, NCCL issues) ## Version Information (as of v0.5.8) | Component | Version | Notes | |-----------|---------|-------| | SGLang | 0.5.8 | Latest stable (2026-01-29) | | sgl-kernel | 0.3.21 | PyPI install for CUDA 12.9 | | mooncake-transfer-engine | 0.3.8.post1 | KV cache transfer (requires nvidia_peermem) | | nixl | 0.9.0 | KV cache transfer (DMA-BUF, recommended) | | nvidia-nccl-cu12 | 2.28.3 | Force reinstall | | nvidia-cudnn-cu12 | 9.16.0.29 | Required for PyTorch 2.9+ | | flashinfer | 0.6.1 | SGLang 0.5.8 requires 0.6.1 (vLLM uses 0.5.3) | | sglang-router | 0.5.8 | PD disaggregation 路由 (pip install sglang-router) | ### What's New in v0.5.8 - **1.5x faster diffusion models** across the board - **Chunked Pipeline Parallelism** for million-token context (near-linear scaling) - **EPD Disaggregation** for Vision-Language Models (elastic encoder scaling) - **GLM4-MoE optimization**: 65% faster TTFT - **New models**: GLM 4.7 Flash, LFM2, Qwen3-VL-Embedding/Reranker, DeepSeek V3.2 NVFP4, FLUX.2-klein-9B ### What's New in v0.5.7 - **Model Gateway v0.3.0** release - **Scalable Pipeline Parallelism** with dynamic chunking for ultra-long contexts - **Encoder Disaggregation** for multi-modal models - **Diffusion improvements**: `--dit-layerwise-offload true` reduces peak VRAM by 30GB - **New models**: Mimo-V2-Flash, Nemotron-Nano-v3, LLaDA 2.0, EAGLE 3 speculative decoding - **Hardware support**: AMD/4090/5090 for diffusion ## Mooncake Transfer Engine Mooncake is required for **prefill-decode disaggregation** mode, which separates prefill and decode phases across different nodes for production deployments. ### Installing Mooncake ```bash pip install --break-system-packages mooncake-transfer-engine==0.3.8.post1 ``` ### Verifying Mooncake ```bash python3 -c "from mooncake.engine import TransferEngine; print('Mooncake OK')" ``` ### When is Mooncake Needed? Mooncake is required when using: - `--disaggregation-mode prefill` or `--disaggregation-mode decode` - Multi-node deployments with KV cache transfer - Production DeepSeek-V3/R1 deployments with prefill-decode separation **Note:** For single-node testing without disaggregation, Mooncake is not required. ## NIXL Transfer Engine (Recommended) NIXL (NVIDIA Inference Xfer Library) is an alternative to Mooncake that uses **DMA-BUF** instead of nvidia_peermem. It's the recommended choice when: - Using NVIDIA Open Kernel Module (nvidia_peermem won't load) - nvidia_peermem fails with "Invalid argument" error - You want a more portable solution that doesn't depend on kernel modules ### Installing NIXL ```bash pip install --break-system-packages nixl==0.9.0 # IMPORTANT: NIXL may downgrade NVIDIA libraries, reinstall correct versions: pip install nvidia-nccl-cu12==2.28.3 --force-reinstall --no-deps pip install nvidia-cudnn-cu12==9.16.0.29 --force-reinstall --no-deps ``` **注意**: NIXL 还会安装 `nvidia-nvshmem-cu12==3.4.5`,这个 pip 包**不会被使用**。 DeepEP 使用的是自编译的 NVSHMEM 3.5.19(带 IBGDA 支持),通过 `unified-env.sh` 中的 `LD_PRELOAD` 加载。 ### Verifying NIXL ```bash python3 -c "import nixl; print('NIXL OK')" ``` ### Using NIXL for Disaggregation Add `--disaggregation-transfer-backend nixl` to your launch command: ```bash python3 -m sglang.launch_server \ --model-path deepseek-ai/DeepSeek-V3 \ --disaggregation-mode prefill \ --disaggregation-transfer-backend nixl \ # Use NIXL instead of Mooncake --tp-size 8 \ ... ``` ### NIXL vs Mooncake | Feature | NIXL | Mooncake | |---------|------|----------| | Memory registration | DMA-BUF (kernel native) | nvidia_peermem (kernel module) | | Transport | UCX (TCP/RDMA/SHM) | RDMA or TCP | | Kernel module required | No | nvidia_peermem (may fail) | | Open Kernel Module compatible | Yes | No (fails to load) | | Recommended for | NVIDIA Open driver, B200 | Legacy systems with nvidia_peermem | **Recommendation:** Use NIXL for new deployments, especially on systems with NVIDIA Open Kernel Module. ## Installation Workflow ### Pre-requisites (Ubuntu 24.04) Ubuntu 24.04 doesn't include pip by default. Install it first: ```bash sudo apt-get update sudo apt-get install -y python3-pip ``` ### Step 1: Environment Setup To set up the environment, ensure CUDA is properly configured: ```bash export CUDA_HOME=/usr/local/cuda export PATH=$CUDA_HOME/bin:$PATH export BUILD_TYPE=blackwell # or "all" for general, "hopper" for H100 ``` ### Step 2: Clone and Install To install SGLang from source: ```bash mkdir -p /sgl-workspace && cd /sgl-workspace # Clone specific version git clone -b v0.5.8 --depth 1 https://github.com/sgl-project/sglang.git cd sglang # Install sgl-kernel first (for CUDA 12.9) pip install sgl-kernel==0.3.21 # Install SGLang with blackwell support pip install -e "python[blackwell]" --extra-index-url https://download.pytorch.org/whl/cu129 ``` ### Step 3: Install Additional Dependencies To install required NVIDIA libraries and NIXL: ```bash # NVIDIA libraries (required) pip install nvidia-nccl-cu12==2.28.3 --force-reinstall --no-deps pip install nvidia-cudnn-cu12==9.16.0.29 --force-reinstall --no-deps # NIXL for KV cache transfer (RECOMMENDED for disaggregation mode) pip install --break-system-packages nixl==0.9.0 # Re-install NVIDIA libs after NIXL (NIXL may downgrade them) pip install nvidia-nccl-cu12==2.28.3 --force-reinstall --no-deps pip install nvidia-cudnn-cu12==9.16.0.29 --force-reinstall --no-deps ``` > ⚠️ **Important**: NIXL is required for prefill-decode disaggregation mode. If you skip NIXL, you'll need nvidia_peermem kernel module (often fails on NVIDIA Open driver). ### Step 4: Configure LD_LIBRARY_PATH To fix library loading issues, run `scripts/setup_env.sh` or manually set: ```bash # Collect all nvidia pip package lib paths NVIDIA_LIB_PATHS="" for d in /usr/local/lib/python3.12/dist-packages/nvidia/*/lib; do [ -d "$d" ] && NVIDIA_LIB_PATHS="${d}:${NVIDIA_LIB_PATHS}" done for d in $HOME/.local/lib/python3.12/site-packages/nvidia/*/lib; do [ -d "$d" ] && NVIDIA_LIB_PATHS="${d}:${NVIDIA_LIB_PATHS}" done export LD_LIBRARY_PATH=${CUDA_HOME}/lib64:${NVIDIA_LIB_PATHS}${LD_LIBRARY_PATH} ``` ## Common Errors and Fixes ### Error: DeepSeek-V3 FP8 outputs garbage on Blackwell (B200) **Symptom:** ``` Model outputs garbage characters like "################" or random symbols ``` **Log Warning:** ``` WARNING model_config.py:872: DeepGemm is enabled but the scale_fmt of checkpoint is not ue8m0. This might cause accuracy degradation on Blackwell. ``` **Diagnosis:** The DeepSeek-V3 official FP8 checkpoint uses `mscale` format which is incompatible with DeepGEMM on Blackwell (B200) GPUs. DeepGEMM expects `ue8m0` scale format. **Workaround:** Use **vLLM** instead of SGLang for DeepSeek-V3 on Blackwell: ```bash # vLLM handles the FP8 scale format correctly vllm serve deepseek-ai/DeepSeek-V3 --tensor-parallel-size 8 --port 8000 --trust-remote-code ``` **Status:** Known issue as of SGLang 0.5.8 on Blackwell GPUs. Works correctly on H100/A100. | Framework | DeepSeek-V3 FP8 on Blackwell | |-----------|------------------------------| | SGLang 0.5.8 | ❌ Garbage output | | vLLM 0.14.1 | ✅ Works correctly | --- ### Error: Cannot uninstall typing_extensions (Ubuntu 24.04) **Symptom:** ``` ERROR: Cannot uninstall typing_extensions 4.10.0, RECORD file not found. Hint: The package was installed by debian. ``` **Diagnosis:** Ubuntu 24.04 installs `typing_extensions` as a system package managed by apt. **Fix:** ```bash pip install -e "python[blackwell]" --ignore-installed typing_extensions --break-system-packages ``` --- ### Error: sgl-kernel ABI mismatch (PyTorch user/system conflict) **Symptom:** ``` ImportError: .../sgl_kernel/sm100/common_ops.abi3.so: undefined symbol: _ZN3c104cuda29c10_cuda_check_implementationEiPKcS2_ib ``` **Diagnosis:** PyTorch installed in both user (`~/.local/lib/python3.12/site-packages/`) and system (`/usr/local/lib/python3.12/dist-packages/`) directories with different versions. **Check:** ```bash pip3 show torch | grep -E "Version|Location" ls /usr/local/lib/python3.12/dist-packages/ | grep torch ls ~/.local/lib/python3.12/site-packages/ | grep torch ``` **Fix:** Remove user-installed torch to use system version: ```bash pip3 uninstall torch torchvision torchaudio -y --break-system-packages python3 -c "import torch; print(torch.__version__)" # Should show system version ``` --- ### Error: libcudnn.so.9 not found **Symptom:** ``` ImportError: libcudnn.so.9: cannot open shared object file ``` **Fix:** ```bash pip install nvidia-cudnn-cu12==9.16.0.29 --force-reinstall --no-deps # Then set LD_LIBRARY_PATH as described above ``` ### Error: libcusparseLt.so.0 not found **Symptom:** ``` ImportError: libcusparseLt.so.0: cannot open shared object file ``` **Fix:** ```bash pip install nvidia-cusparselt-cu12 # Then set LD_LIBRARY_PATH as described above ``` ### Error: assert self.total_num_heads % tp_size == 0 **Symptom:** ``` AssertionError: assert self.total_num_heads % tp_size == 0 ``` **Diagnosis:** The model's attention head count is not divisible by the tensor parallelism size. **Fix:** Choose a `--tp` value that divides the model's attention head count: | Model | Attention Heads | Valid TP Values | |-------|-----------------|-----------------| | Qwen2.5-7B | 28 | 1, 2, 4, 7, 14 | | Qwen2.5-72B | 64 | 1, 2, 4, 8, 16, 32 | | Llama-3-8B | 32 | 1, 2, 4, 8, 16, 32 | | Llama-3-70B | 64 | 1, 2, 4, 8, 16, 32 | | DeepSeek-R1 | 128 | 1, 2, 4, 8, 16, 32, 64 | To find the attention head count for any model: ```bash python3 -c "from transformers import AutoConfig; c = AutoConfig.from_pretrained('MODEL_NAME'); print(f'Attention heads: {c.num_attention_heads}')" ``` ### Error: NCCL errors or timeouts **Fix:** ```bash pip install nvidia-nccl-cu12==2.28.3 --force-reinstall --no-deps ``` ### Error: sgl-kernel version mismatch **Symptom:** SGLang installs an older sgl-kernel version than expected. **Note:** As of v0.5.6.post2, SGLang's dependencies pin sgl-kernel to 0.3.19, so even if you pre-install 0.3.21, it will be downgraded during SGLang installation. This is expected behavior and 0.3.19 works correctly. **If you need a specific version:** Install sgl-kernel AFTER SGLang: ```bash pip install -e "python[blackwell]" ... pip install sgl-kernel==0.3.21 --force-reinstall --no-deps # if needed ``` ### Error: sgl-kernel ABI incompatibility (undefined symbol) **Symptom:** ``` ImportError: .../sgl_kernel/sm100/common_ops.abi3.so: undefined symbol: _ZN3c104cuda29c10_cuda_check_implementationEiPKcS2_ib
عرض على GitHub
ملف SKILL.md هذا كبير جدا، لذلك يعرض SkillsMP القسم الاول فقط هنا. عرض على GitHub