Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jit_kernel module
原文の言語: 英語
メニュー
SkillsMP は FutureMLS-Lab/OSCAR から 11 件の skill を収集しています。skill を開くとソースと詳細を確認できます。
収集済み skill 11 件中 11 件を表示しています。
Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jit_kernel module
原文の言語: 英語
Step-by-step tutorial for adding a heavyweight AOT CUDA/C++ kernel to sgl-kernel (including tests & benchmarks)
原文の言語: 英語
Guide to SGLang CI workflow orchestration — stage ordering, fast-fail, gating, partitioning, execution modes, and debugging CI failures. Use when modifying CI workflows, adding stages, debugging CI pipeline issues, or understanding how tests are dispatched…
原文の言語: 英語
Generate an e2e profiling trace of an SGLang server run. Launches a server, validates accuracy, captures a Chrome-compatible trace, and returns the profile path.
原文の言語: 英語
Run SGLang auto benchmark searches with tiered server-flag sweeps, canonical dataset preparation, ShareGPT auto-download, custom-data conversion/validation, SLA or fixed-QPS benchmarking, CSV export, and optional second-stage speculative/EAGLE tuning. Use…
原文の言語: 英語
Compact SGLang torch-profiler triage skill. Use when Codex should inspect an existing `trace.json(.gz)` or profile directory, trigger `sglang.profiler` against a live server, and return one compact report with kernel, overlap-opportunity, and fuse-pattern…
原文の言語: 英語
Guide for writing SGLang CI/UT tests. Covers CustomTestCase, CI registration, server fixtures, model selection, mock testing, and test placement. Always read test/README.md for the full CI layout, how to run tests, and extra tips. Use when creating new tests,…
原文の言語: 英語
Use when adding a new diffusion model or Diffusers pipeline to SGLang.
原文の言語: 英語
Use when optimizing an existing SGLang diffusion kernel with AKO4ALL, including AKO4ALL repo hygiene, custom microbench setup, ncu-guided iteration, and end-to-end denoise validation. Also use when a sibling AKO4ALL repo must be cloned or refreshed before…
原文の言語: 英語
Use when benchmarking denoise latency or profiling a diffusion bottleneck in SGLang.
原文の言語: 英語
Use when choosing the fastest SGLang Diffusion flags for a model, GPU, and VRAM budget.
原文の言語: 英語