Skip to main content

vllm-project/vllm-skills

SkillsMP has collected 6 skills from vllm-project/vllm-skills. Open a skill to review its source and details.

Latest recorded source activity
SkillsMP catalog refreshed
skills collected
6
GitHub stars
95
GitHub forks
27

Skills in this repository

Showing 6 of 6 collected skills.

occupation
Data Scientists
description

Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics. Use when the user wants to quickly test vLLM serving performance without downloading…

updated
occupation
Data Scientists
description

Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve. Supports multiple datasets (random, sharegpt, sonnet, HF), backends (openai, openai-chat, vllm-pooling, embeddings), throughput/latency testing with request-rate control, and result…

updated
occupation
Network & Computer Systems Administrators
description

Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

updated
occupation
Network & Computer Systems Administrators
description

Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint. Use this skill whenever the user wants to deploy, run, or serve vLLM on a Kubernetes cluster, including creating deployments, services, checking existing…

updated
occupation
Software Developers
description

Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.

updated
occupation
Software Developers
description

This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns. Use when the user asks to benchmark prefix caching hit rate, caching efficiency, or…

updated
Showing 6 of 6 collected skills.