| name | vllm-serving-setup |
| description | Design, deploy, and tune vLLM v0.18.2 inference serving on EKS with PagedAttention v2, Multi-LoRA, FP8 KV Cache, Chunked Prefill, and Continuous Batching. Produces Helm values.yaml, PodMonitor, HPA, and kubectl validation steps for production agentic workloads. |
| argument-hint | [model name, target QPS/latency, GPU budget] |
| user-invocable | true |
| model | claude-sonnet-4-6 |
| allowed-tools | Read,Write,Edit,Bash,Grep,Glob,mcp__eks,mcp__aws-documentation,mcp__aws-pricing,mcp__cloudwatch,mcp__prometheus |
When to Use
- ํน์ ๋ชจ๋ธ(์: Llama 3.3 70B, Qwen3 32B, DeepSeek-V3)์ vLLM ์ผ๋ก EKS ์ ๋ฐฐํฌํ ๋
- ๊ธฐ์กด vLLM ๋ฐฐํฌ์ ์ฒ๋ฆฌ๋ยท์ง์ฐ์ ํ๋ํ ๋
- Multi-LoRA ์ด๋ํฐ๋ฅผ ํ๋์ base ๋ชจ๋ธ์ ์ฌ๋ฆด ๋
- FP8 KV Cache, Chunked Prefill ๋ฑ v0.18.2 ์ ๊ธฐ๋ฅ์ ํ์ฑํํ ๋
When NOT to Use
- ๋ถ์ฐ ์ถ๋ก (Disaggregated Prefill/Decode)์ด ํ์ํ ๋๊ท๋ชจ ํธ๋ํฝ โ
llm-d ์คํฌ ์ฌ์ฉ ๊ถ์ฅ
- ์ฌ์ ํ์ต(training) ํ์ดํ๋ผ์ธ โ SageMaker / KubeRay / SkyPilot ์์ญ
- MoE ์ ์ฉ ์ต์ ํ๊ฐ ํ์ํ ๊ฒฝ์ฐ โ Expert Parallel ์ ์ฉ ์ค์ ์ฐธ์กฐ
Preconditions
- EKS ํด๋ฌ์คํฐ์ GPU NodePool, GPU Operator, DCGM ์ด ์ ์ ๋์ (
agentic-eks-bootstrap ์๋ฃ)
- Hugging Face ์ก์ธ์ค ํ ํฐ(Secrets Manager ์ ์ฅ) ๋๋ S3 ๋ชจ๋ธ ๊ฐ์ค์น ๋ณต์ฌ ์๋ฃ
- Prometheus + OTel Collector ๊ฐ ๋ฐฐํฌ๋์ด ์์
Procedure
Step 1. GPU ๋ฉ๋ชจ๋ฆฌ ์ฌ์ด์ง
ํ์ GPU ๋ฉ๋ชจ๋ฆฌ = ๋ชจ๋ธ ๊ฐ์ค์น + ๋นtorch ๋ฉ๋ชจ๋ฆฌ
+ PyTorch ํ์ฑํ ํผํฌ ๋ฉ๋ชจ๋ฆฌ
+ (๋ฐฐ์น๋น KV ์บ์ ร ๋ฐฐ์น ํฌ๊ธฐ)
- Llama 3.3 70B FP16 โ ๊ฐ์ค์น 140GB + KV ์บ์ ~40GB + ์ค๋ฒํค๋ 20GB โ 200GB
- ๋จ์ผ H100 80GB ๋ถ๊ฐ โ TP=4 (GPU๋น 50GB)
- INT4 ์์ํ ์ 35GB โ ๋จ์ผ A100 80GB ๋๋ H100 ๊ฐ๋ฅ
Step 2. ๋ณ๋ ฌํ ์ ๋ต ์ ์
- TP (Tensor Parallel): ๋์ผ ๋
ธ๋ ๋ด GPU ๊ฐ layer ํ๋ผ๋ฏธํฐ ๋ถ์ฐ
- PP (Pipeline Parallel): ๋ ์ด์ด ๊ทธ๋ฃน์ ๋
ธ๋ ๊ฐ ๋ถ์ฐ (๋ฉํฐ ๋
ธ๋)
- EP (Expert Parallel): MoE ๋ชจ๋ธ ์ ์ฉ
- DP (Data Parallel): replica ํ์ฅ, HPA ์ ๊ฒฐํฉ
Step 3. Helm values ์์ฑ
model:
name: meta-llama/Llama-3.3-70B-Instruct
dtype: bfloat16
vllm:
image: vllm/vllm-openai:v0.18.2
extraArgs:
- --tensor-parallel-size=4
- --gpu-memory-utilization=0.92
- --max-model-len=32768
- --enable-chunked-prefill
- --enable-prefix-caching
- --kv-cache-dtype=fp8
- --enable-lora
- --max-loras=8
- --otlp-traces-endpoint=http://otel-collector.observability.svc:4317
resources:
limits:
nvidia.com/gpu: 4
nodeSelector:
karpenter.sh/nodepool: gpu-inference-pool
tolerations:
- key: nvidia.com/gpu
operator: Exists
Step 4. ๋ฐฐํฌ ๋ฐ ๊ฒ์ฆ
helm upgrade --install vllm-llama3 vllm-project/vllm \
--namespace inference --create-namespace \
-f values.yaml --version 0.18.2
kubectl -n inference rollout status deployment/vllm-llama3 --timeout=10m
kubectl -n inference port-forward svc/vllm-llama3 8000:8000 &
curl -s http://localhost:8000/v1/models | jq
curl -s http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"meta-llama/Llama-3.3-70B-Instruct","messages":[{"role":"user","content":"hello"}]}'
Step 5. HPA + PodMonitor
- HPA:
vllm:num_requests_running ๋๋ DCGM_FI_DEV_GPU_UTIL ๊ธฐ๋ฐ
- PodMonitor:
/metrics ์๋ํฌ์ธํธ (prometheus_client)
- Langfuse ๋ก OTel trace ์ ์ก๋๋์ง ํ์ธ
Step 6. ๋ฒค์น๋งํฌ
python benchmarks/benchmark_serving.py \
--backend vllm \
--model meta-llama/Llama-3.3-70B-Instruct \
--num-prompts 200 \
--request-rate 10
- p50/p95/p99 latency, throughput (tokens/s), GPU util ์บก์ฒ
Good Examples
- Llama 3.3 70B: TP=4, FP8 KV Cache, Chunked Prefill on โ 30% throughput ๊ฐ์
- Qwen3 32B: TP=2, Prefix Caching ํ์ฑ โ code ๋ฒค์น๋งํฌ 400%+ ๊ฐ์
- Multi-LoRA 16๊ฐ: ๋จ์ผ base + ์ด๋ํฐ hot swap, ๊ณ ๊ฐ๋ณ tenant ๋ถ๋ฆฌ
Bad Examples (๊ธ์ง)
gpu-memory-utilization=0.98 โ OOM ์ํ, 0.92 ์ดํ ๊ถ์ฅ
--max-model-len ์์ ์ถ์ ์์ด ๊ธด ํ๋กฌํํธ ํ์ฉ โ KV ์บ์ ํญ์ฃผ
- TP ์ค์ ์ด GPU ๊ฐ์์ ๋ถ์ผ์น (์: TP=4 ์ธ๋ฐ request 3 GPU) โ ๋ถํ
์คํจ
- vLLM v0.6.x ๊ตฌ๋ฒ์ ์ฌ์ฉ โ Chunked Prefill, FP8 KV ๋ฏธ์ง์
References