Skip to main content

vllm-bench-serve

Interactive online benchmark orchestrator for vLLM inference services using `vllm bench serve`. Supports single benchmarks, multi-case batch execution with result aggregation, and auto-optimization to find optimal concurrency/throughput under latency SLO constraints (TTFT, TPOT, P99, success rate). Use this skill whenever the user wants to benchmark, stress test, or measure performance of a running LLM/multimodal/embedding inference service — even if they don't say "vllm bench serve" explicitly. Do NOT use for offline inference throughput, service deployment/startup, profiling/tracing, health checks only, or analyzing existing benchmark results without running new tests.

Zur Installation springen

Quellinformationen

Repository
ascend-ai-coding/awesome-ascend-skills
Letzte Quellaktivität
18. Mai 2026 um 12:39
Erkannte Sprache von SKILL.md
Englisch
Sterne
166
Forks
56

Installationsoptionen

Standardmäßig ist der Prompt ausgewählt, der zuerst die Quelle prüft. Sie können zu einem direkten Befehl wechseln oder eine lokale Kopie herunterladen.

Quelldateien prüfen

Lesen Sie SKILL.md und alle von SkillsMP angezeigten Begleitdateien, bevor Sie sich für eine Installation entscheiden.