Skip to main content

vllm-project/vllm-skills

SkillsMP는 vllm-project/vllm-skills에서 6개의 skill을 수집했습니다. skill을 열어 소스와 세부 정보를 확인하세요.

최근 기록된 소스 활동
SkillsMP 카탈로그 업데이트
수집된 skills
6
GitHub 스타
95
GitHub 포크
27

이 저장소의 skills

수집된 skill 6개 중 6개를 표시합니다.

직업 분류
데이터 과학자
설명

Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics. Use when the user wants to quickly test vLLM serving performance without downloading…

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve. Supports multiple datasets (random, sharegpt, sonnet, HF), backends (openai, openai-chat, vllm-pooling, embeddings), throughput/latency testing with request-rate control, and result…

원문 언어: 영어

업데이트
직업 분류
네트워크·컴퓨터 시스템 관리자
설명

Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

원문 언어: 영어

업데이트
직업 분류
네트워크·컴퓨터 시스템 관리자
설명

Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint. Use this skill whenever the user wants to deploy, run, or serve vLLM on a Kubernetes cluster, including creating deployments, services, checking existing…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns. Use when the user asks to benchmark prefix caching hit rate, caching efficiency, or…

원문 언어: 영어

업데이트
수집된 skill 6개 중 6개를 표시합니다.