Skip to main content

bench-open-model

Benchmark any open-weight vLLM-servable Hugging Face model on AWS GPU hardware. Sizes the model's VRAM needs, picks the right EC2 GPU instance type and region, provisions it, serves the model on vLLM, runs a cache-honest concurrency sweep, writes a report, and tears everything down. Use when asked to benchmark / test / evaluate the throughput, latency, or serving capacity of an open-weight model on AWS, pick a GPU instance type for a model, size VRAM for a model, or measure tokens-per-second or requests-per-day. Covers text LLMs and multimodal / vision / OCR models.

Jump to install

Source facts

Repository
aws-samples/sample-gpu-open-weight-bench
Last source activity
August 6, 2026 at 08:50
Detected SKILL.md language
English
Stars
1
Forks
0

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.