Skip to main content

vllm-project/vllm-skills

SkillsMP 已收集 vllm-project/vllm-skills 中的 6 个 Skill。打开任一 Skill 可查看来源和详情。

最近记录的来源活动
SkillsMP 收录数据更新
已收集 skills
6
GitHub 星标
95
GitHub Forks
27

这个仓库中的 skills

已展示 6 / 6 个已收集 Skill。

职业分类
数据科学家
描述

Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics. Use when the user wants to quickly test vLLM serving performance without downloading…

原文语言:英语

更新
职业分类
数据科学家
描述

Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve. Supports multiple datasets (random, sharegpt, sonnet, HF), backends (openai, openai-chat, vllm-pooling, embeddings), throughput/latency testing with request-rate control, and result…

原文语言:英语

更新
职业分类
网络与计算机系统管理员
描述

Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

原文语言:英语

更新
职业分类
网络与计算机系统管理员
描述

Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint. Use this skill whenever the user wants to deploy, run, or serve vLLM on a Kubernetes cluster, including creating deployments, services, checking existing…

原文语言:英语

更新
职业分类
软件开发工程师
描述

Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.

原文语言:英语

更新
职业分类
软件开发工程师
描述

This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns. Use when the user asks to benchmark prefix caching hit rate, caching efficiency, or…

原文语言:英语

更新
已展示 6 / 6 个已收集 Skill。