用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/kunpengcompute/sglang --skill sglang-diffusion-benchmark-profile命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
Use when quantizing a diffusion DiT with NVIDIA ModelOpt and making the resulting FP8 or NVFP4 checkpoint loadable, verifiable, and benchmarkable in SGLang Diffusion.
Use when choosing the fastest SGLang Diffusion flags for a model, GPU, and VRAM budget.
基于 SOC 职业分类
正在显示 SKILL.md
| name | sglang-diffusion-benchmark-profile |
| description | Use when benchmarking denoise latency or profiling a diffusion bottleneck in SGLang. |
Use this skill when measuring denoise performance, finding the slow op, checking whether an existing fast path can solve it, or verifying that a hotspot is real before any kernel work in sglang.multimodal_gen.
This skill is diagnosis-first. It owns:
torch.profiler trace capture and quick hotspot rankingThis skill does not own low-level kernel authoring or standalone Nsight workflows.
Before running any benchmark, profiler, or kernel-validation command:
scripts/diffusion_skill_env.py to derive the repo root from sglang.__file__HF_TOKEN before using gated Hugging Face models such as black-forest-labs/FLUX.*FLASHINFER_DISABLE_VERSION_CHECK=1All diffusion benchmark and profiling results owned by this skill must come from the native SGLang diffusion backend.
Treat any of the following as a hard stop condition:
Falling back to diffusers backendUsing diffusers backendLoaded diffusers pipelineIf any benchmark, perf-dump, or torch.profiler command prints one of those signals:
torch.profiler workflow; uses checked-in nightly-aligned presets plus skill-only stress recipes such as LTX-2.3 one-stage/two-stage, HunyuanVideo, MOVA, and HeliosQK norm + RoPE, distributed overlap patterns, and open optimization PRs before proposing new codesglang.__file__, write-access probe, benchmark/profile output directories, idle GPU selectionsglang generate; pins --backend=sglang, supports --no-torch-compile, and saves perf dumps by label for compare_perf.pyBefore calling a diffusion hotspot "new", first classify it with existing-fast-paths.md.
Always rule out these existing families first:
QK norm + RoPEtorch.compile compute / communication reorderIf the user explicitly requires torch.compile to stay off, do not use the
default benchmark preset invocation unchanged. Either pass the checked-in
benchmark helper its no-compile switch or run the equivalent manual command
without --enable-torch-compile.
For FLUX-family manual profiling runs with a quantized transformer override:
sglang generate directly--transformer-path <dir>--prompt-path <file> when also fixing --output-file-name--model-path plus HF_HUB_OFFLINE=1--profile changes latency substantially; use the non-profile perf dump for the real before/after benchmark claim