Skip to main content

sgl-project/sglang

O SkillsMP coletou 28 skills de sgl-project/sglang. Abra uma skill para revisar a origem e os detalhes.

Última atividade de origem registrada
Catálogo do SkillsMP atualizado
skills coletadas
28
Estrelas no GitHub
32.000
Forks no GitHub
7.972

Mostrando 28 de 28 skills coletadas.

ocupação
Desenvolvedores de software
descrição

How SGLang's runtime configuration and process-global state are organized (RuntimeContext tiers, publish + namespace config bags, the pristine ServerArgs seed, override entry points, resource/stream/buffer leases, per-forward flags), the CI guardrails that…

Idioma do texto original: inglês

atualizado
ocupação
Analistas de garantia de qualidade de software e testadores
descrição

Start and persistently pursue a goal to babysit an SGLang pull request until selected GitHub Actions workflows pass on the latest PR head. Use when asked to monitor, babysit, retry, or fix PR CI for lint.yml, pr-test.yml, pr-test-extra.yml, AMD, or other…

Idioma do texto original: inglês

atualizado
ocupação
Desenvolvedores de software
descrição

Add a new model to the SGLang Cookbook (docs/, Mintlify), config-driven format — instantiate the model-agnostic template into a per-model config (+ benchmarks) JSX under src/snippets/configs/, an MDX page, the docs.json nav entry, NEW-tag hygiene, and the…

Idioma do texto original: inglês

atualizado
ocupação
Desenvolvedores de software
descrição

Use when benchmarking denoise latency or profiling a diffusion bottleneck in SGLang.

Idioma do texto original: inglês

atualizado
ocupação
Desenvolvedores de software
descrição

Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jit_kernel module

Idioma do texto original: inglês

atualizado
ocupação
Analistas de garantia de qualidade de software e testadores
descrição

Guide for writing SGLang CI/UT tests. Covers CustomTestCase, CI registration, server fixtures, model selection, mock testing, and test placement. Always read test/README.md for the full CI layout, how to run tests, and extra tips. Use when creating new tests,…

Idioma do texto original: inglês

atualizado
ocupação
Analistas de garantia de qualidade de software e testadores
descrição

Write, calibrate, and debug the prefill-vs-decode logprob (KL) consistency tests in sglang -- the two independent conditions a zero requires (every operator batch-invariant, and the two paths computing the same function), which helper separates them, how to…

Idioma do texto original: inglês

atualizado
ocupação
Desenvolvedores de software
descrição

Use when adding a new diffusion model or Diffusers pipeline to SGLang.

Idioma do texto original: inglês

atualizado
ocupação
Desenvolvedores de software
descrição

Use when quantizing a diffusion DiT with NVIDIA ModelOpt and making the resulting FP8 or NVFP4 checkpoint loadable, verifiable, and benchmarkable in SGLang Diffusion.

Idioma do texto original: inglês

atualizado
ocupação
Desenvolvedores de software
descrição

Use when choosing the fastest SGLang Diffusion flags for a model, GPU, and VRAM budget.

Idioma do texto original: inglês

atualizado
ocupação
Desenvolvedores de software
descrição

Migrate a legacy-template SGLang cookbook page (monolithic per-model generator under docs/src/snippets/autoregressive/) onto the config-driven template (shared _deployment.jsx / _playground.jsx engines + per-model config). Use when asked to migrate, convert,…

Idioma do texto original: inglês

atualizado
ocupação
Analistas de garantia de qualidade de software e testadores
descrição

Review a pull request against the SGLang Cookbook (docs/, Mintlify) contribution checklist — the config-driven format (per-model config + benchmarks JSX consumed by the shared _deployment.jsx / _playground.jsx engines). Run with /cookbook-review-pr <PR…

Idioma do texto original: inglês

atualizado
ocupação
Desenvolvedores de software
descrição

Unified LLM torch-profiler triage skill for `sglang`, `vllm`, `TensorRT-LLM`, and `TokenSpeed`. Use it to inspect an existing `trace.json(.gz)` or profile directory, or to drive live profiling against a running server when supported and return one three-table…

Idioma do texto original: inglês

atualizado
ocupação
Desenvolvedores de software
descrição

Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and serving config. Use when a user asks what ratio to set, why…

Idioma do texto original: inglês

atualizado
ocupação
Desenvolvedores de software
descrição

Step-by-step tutorial for adding a heavyweight AOT CUDA/C++ kernel to sgl-kernel (including tests & benchmarks)

Idioma do texto original: inglês

atualizado
ocupação
Desenvolvedores de software
descrição

Make mechanical refactoring (file splits, function moves, module extractions, renames) machine-checkable instead of eyeballed. Reproduce a relocation commit byte-for-byte from faithful primitives, and split an extraction into a verifiable prepare + move +…

Idioma do texto original: inglês

atualizado
ocupação
Desenvolvedores de software
descrição

Code style for SGLang large classes `Scheduler`, `TokenizerManager`, and `ModelRunner`: frozen-code conventions and `__init__` orchestration style. Use when modifying any of these three classes or reviewing changes to them.

Idioma do texto original: inglês

atualizado
ocupação
Desenvolvedores de software
descrição

Guide to SGLang CI workflow orchestration — stage ordering, fast-fail, gating, partitioning, execution modes, and debugging CI failures. Use when modifying CI workflows, adding stages, debugging CI pipeline issues, or understanding how tests are dispatched…

Idioma do texto original: inglês

atualizado
ocupação
Desenvolvedores de software
descrição

Conventions for SGLang environment variables — where to define, how to access, how to name, and how to deprecate. Use when adding, renaming, or reviewing any `SGLANG_*` environment variable (or migrating a legacy `SGL_*` alias), or when touching…

Idioma do texto original: inglês

atualizado
ocupação
Desenvolvedores de software
descrição

Requirements for the SGLang scripted runtime, chiefly when to add (vs not add) a harness API. Use for anything related to the scripted runtime.

Idioma do texto original: inglês

atualizado
ocupação
Desenvolvedores de software
descrição

Naming conventions for SGLang speculative decoding identifiers. Use when adding, renaming, or reviewing identifiers in speculative decoding code — anything under `python/sglang/srt/speculative/`, related attention backends, scheduler accumulators, IPC fields,…

Idioma do texto original: inglês

atualizado
ocupação
Desenvolvedores de software
descrição

Clean up noisy startup warnings and spurious prints in SGLang server logs. Use when users ask to clean up unwanted warnings, deprecation messages, or third-party noise in the server startup output.

Idioma do texto original: inglês

atualizado
ocupação
Desenvolvedores de software
descrição

Trigger the bot-cherry-pick workflow for a batch of merged PRs onto a release branch and monitor each run to completion. Use when an SGLang release manager asks to cherry-pick a list of PRs to a release branch.

Idioma do texto original: inglês

atualizado
ocupação
Administradores de redes e sistemas de computador
descrição

Replay-first debug flow for SGLang serving problems. Use when a live or recent server shows health-check failures, latency or throughput regressions, queue growth, timeouts, distributed stalls, crash dumps, wrong outputs after deploys, or PD/EP/HiCache…

Idioma do texto original: inglês

atualizado
ocupação
Desenvolvedores de software
descrição

Generate an e2e profiling trace of an SGLang server run. Launches a server, validates accuracy, captures a Chrome-compatible trace, and returns the profile path.

Idioma do texto original: inglês

atualizado
ocupação
Analistas de garantia de qualidade de software e testadores
descrição

Investigate consistently failing SGLang CI tests by extracting the failure signature from scheduled or rerun workflows, bisecting the passing/failing commit window, checking runner or hardware specificity, and optionally reproducing on a remote GPU host.

Idioma do texto original: inglês

atualizado
ocupação
Desenvolvedores de software
descrição

Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP). Covers identifying hang locations via py-spy/watchdog/cuda coredump, per-rank logging to find state divergence, binary-search methodology for locating the first diverge point, and fix…

Idioma do texto original: inglês

atualizado
ocupação
Desenvolvedores de software
descrição

Call this skill when you need to debug CUDA crashes in SGLang using kernel API logging

Idioma do texto original: inglês

atualizado
Mostrando 28 de 28 skills coletadas.