Skip to main content

benchmark-vllm

Measures vLLM serving performance on the zentorch/ZenDNN CPU backend on AMD EPYC. Brings up a multi-instance stack -- N cpuset-pinned `vllm serve` containers, one core slice each, behind an NGINX round-robin load balancer -- and drives it with GuideLLM through the redhat-et/vllm-cpu-perf-eval ansible automation in external-endpoint mode. Sizes the instance layout from the machine's real topology (SMT siblings kept inside one instance), derives the server context window from the selected workloads, and sweeps a list of request rates against one warm stack, so throughput numbers (high concurrency, 32/64+) and latency numbers (low concurrency, 1/2/4/8) come from an identical stack and stay comparable. Can also run a native (non-zentorch) A/B leg from the same image with zentorch uninstalled, so a zentorch-vs-native delta is not confounded by a different vLLM commit, torch build or compiler flags. Runs the perf-eval workload set (chat, rag, code, summarization, reasoning, variable-length and context-scaling profi

الانتقال إلى التثبيت

معلومات المصدر

المستودع
amd/ZenDNN-pytorch-plugin
آخر نشاط في المصدر
٥ أغسطس ٢٠٢٦ في ١٢:٠١
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
٣٦
التفرعات
١٠

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.