Skip to main content

sglang

Comprehensive reference documentation and skill for SGLang - a high-performance serving framework for large language models and multimodal models. Covers SGLang architecture, ServerArgs configuration, OpenAI-compatible API server, native API, offline engine API, attention backends (FlashInfer, FlashAttention, Triton, FlashMLA, cutlass_mla), KV cache management with RadixAttention, paged attention, sampling and decoding (structured outputs, constrained decoding, speculative decoding with EAGLE/ngram/DFlash), distributed inference (tensor/pipeline/expert/data parallelism), PD disaggregation, EPD disaggregation, quantization (FP8, FP4/MXFP4, GPTQ, AWQ, INT4/INT8, Marlin, bitsandbytes, GGUF, modelopt), multi-LoRA batching, multimodal processing (image/audio/video), CUDA graphs, torch.compile, piecewise CUDA graphs, sgl-kernel, sgl-model-gateway (Rust), HiCache, HiSparse, RL/post-training support, checkpoint engine, diffusion models, observability, profiling, supported model architectures (200+ models), hardware p

الانتقال إلى التثبيت

معلومات المصدر

المستودع
jstzwj/ai-infra-plugins
آخر نشاط في المصدر
٧ مايو ٢٠٢٦ في ١١:٢١
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
٤
التفرعات
٠

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.