Skip to main content

vllm-bench-serve

Interactive online benchmark orchestrator for vLLM inference services using `vllm bench serve`. Supports single benchmarks, multi-case batch execution with result aggregation, and auto-optimization to find optimal concurrency/throughput under latency SLO constraints (TTFT, TPOT, P99, success rate). Use this skill whenever the user wants to benchmark, stress test, or measure performance of a running LLM/multimodal/embedding inference service — even if they don't say "vllm bench serve" explicitly. Do NOT use for offline inference throughput, service deployment/startup, profiling/tracing, health checks only, or analyzing existing benchmark results without running new tests.

الانتقال إلى التثبيت

معلومات المصدر

المستودع
ascend-ai-coding/awesome-ascend-skills
آخر نشاط في المصدر
١٨ مايو ٢٠٢٦ في ١٢:٣٩
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
١٦٦
التفرعات
٥٦

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.