| name | agentic-eks-bootstrap |
| description | Bootstrap an AWS EKS cluster optimized for Agentic AI workloads โ Karpenter v1.2+ GPU node pools, EKS Auto Mode, Kubernetes 1.32+ with DRA 1.35 GA, VPC CNI, GPU Operator, and baseline observability. Use when starting a new EKS cluster that will host vLLM, Inference Gateway, Langfuse, or Kagent. |
| argument-hint | [cluster-name, region, expected GPU workload profile] |
| user-invocable | true |
| model | claude-sonnet-4-6 |
| allowed-tools | Read,Write,Edit,Bash,Grep,Glob,mcp__eks,mcp__aws-documentation,mcp__aws-iac,mcp__aws-pricing,mcp__well-architected-security |
When to Use
- ์ ๊ท Agentic AI ํ๋ซํผ์ EKS ์์ ๊ตฌ์ถํ๊ธฐ ์์ํ ๋
- ๊ธฐ์กด EKS ํด๋ฌ์คํฐ๋ฅผ GPUยทAgent ์ํฌ๋ก๋์ฉ์ผ๋ก ์ฌ๊ตฌ์ฑํ ๋
platform-architect ๊ฐ "EKS ๊ฒฝ๋ก"๋ฅผ ํ์ ํ ์ดํ ์ค์ ํด๋ฌ์คํฐ ํ๋ก๋น์ ๋ ๋จ๊ณ
When NOT to Use
- Bedrock AgentCore ๋๋ SageMaker Unified Studio ๋ง ์ฌ์ฉํ ์์ ์ผ ๋ โ ์ด ์คํฌ์ ๋ถํ์ํฉ๋๋ค
- ์ด๋ฏธ KarpenterยทGPU Operator ๊ฐ ์ ์ ๋์ํ๋ ๊ธฐ์กด ํด๋ฌ์คํฐ โ
vllm-serving-setup ์ผ๋ก ๋ฐ๋ก ์งํ
- PoC ์์ค์ ๋ก์ปฌ k3d/kind ํ๊ฒฝ โ EKS ์ ์ฉ ๊ธฐ๋ฅ(IRSA, Karpenter EC2 ์๋ ํ๋ก๋น์ ๋)์ด ๋ถํ์
Preconditions
- AWS CLI ๋ฐ
eksctl v0.196+, helm v3.14+, kubectl v1.32+ ์ค์น
- ๊ด๋ฆฌ์ IAM ๊ถํ ๋๋ EKS ํด๋ฌ์คํฐ ์์ฑ ๊ฐ๋ฅํ Role
- ๋์ ๋ฆฌ์ ์์ ์๊ตฌ GPU ์ธ์คํด์ค ์ฟผํฐ ํ์ธ (
p5, g6e, trn2)
Procedure
Step 1. Platform ์๊ตฌ์ฌํญ ์์ง
- ์์ ๋์ ์ฌ์ฉ์/QPS, ๋ชจ๋ธ ํฌ๊ธฐ, SLA(์ง์ฐ) ๊ฐ ํ์ธ
- ํ๊ตญ ๊ธ์ต๊ถ ๊ท์ ๋์ ์ฌ๋ถ(ISMS-P, ์ ์๊ธ์ต๊ฐ๋
๊ท์ ) ํ์ธ
- Private / Hybrid / Public ๋ฐฐํฌ ์คํ์ผ ๊ฒฐ์
Step 2. ํด๋ฌ์คํฐ ์์ฑ (EKS Auto Mode ๊ถ์ฅ)
eksctl create cluster \
--name agentic-prod \
--region ap-northeast-2 \
--version 1.32 \
--auto-mode \
--with-oidc \
--zones ap-northeast-2a,ap-northeast-2c
- Auto Mode ๋ Karpenter, EBS CSI, VPC CNI, CoreDNS ๋ฅผ AWS ๊ฐ ๊ด๋ฆฌํฉ๋๋ค
- ์๋ ๊ด๋ฆฌ๋ฅผ ์ํ๋ฉด
--node-type ์ง์ + Karpenter ๋ณ๋ ์ค์น
Step 3. Karpenter GPU NodePool ์์ฑ
karpenter.sh/v1 API ์ฌ์ฉ, capacity-type ์ on-demand + spot ํผํฉ
nvidia.com/gpu taint ๋ก GPU ๋
ธ๋ ๊ฒฉ๋ฆฌ
consolidation ์ ์ฑ
์ผ๋ก idle GPU ์๋ ํ์
Step 4. NVIDIA GPU Operator ์ค์น
helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
helm install gpu-operator nvidia/gpu-operator \
--namespace gpu-operator --create-namespace \
--version v24.6.2 \
--set driver.enabled=true \
--set toolkit.enabled=true \
--set dcgmExporter.enabled=true
- Kubernetes 1.32+ ์์ DRA 1.35 GA ๋ฅผ ํ์ฉํ๋ ค๋ฉด
--feature-gates=DynamicResourceAllocation=true
Step 5. ๋ฒ ์ด์ค๋ผ์ธ Addon
- AWS Load Balancer Controller (ALB/NLB)
- External Secrets Operator (IRSA + Secrets Manager)
- Prometheus Stack (kube-prometheus-stack)
- Fluent Bit โ CloudWatch Logs
- Cert-Manager (ACME)
Step 6. ๋ณด์ ๋ฒ ์ด์ค๋ผ์ธ
0.0.0.0/0 SG ์คํ ๊ธ์ง, ๋ด๋ถ ALB + Cognito/OIDC ๊ฒฝ์
- IRSA: Karpenter, GPU Operator, Langfuse ๊ฐ๊ฐ ์ ์ฉ Role
- CIS EKS Benchmark ์๋ ์ค์บ (kube-bench)
well-architected-security MCP ๋ก SEC-01~SEC-11 ์ ๊ฒ
Step 7. ๊ฒ์ฆ
kubectl get nodes -L karpenter.sh/nodepool,node.kubernetes.io/instance-type
kubectl -n gpu-operator get pods
kubectl get gatewayclass
kubectl get crd | grep -E 'nodepool|gpu|dra'
Good Examples
- ํ๋ก๋์
:
--version 1.32, Auto Mode on, p5.48xlarge + g6e ํผํฉ NodePool, DCGM + Prometheus
- ํ์ด๋ธ๋ฆฌ๋: Bedrock ๋งค๋์ง๋ ๊ธฐ๋ณธ + EKS burst pool (Spot 70% / On-Demand 30%)
Bad Examples (๊ธ์ง)
- Security Group
0.0.0.0/0 inbound โ ํ์ฌ ์ ์ฑ
์๋ฐ
--version 1.28 ์ดํ โ DRA ๋ฏธ์ง์, EOL ์๋ฐ
- Karpenter v0.x (legacy API) โ v1 migration ํ์
- GPU Operator ์์ด nvidia-device-plugin ๋จ๋
์ฌ์ฉ โ DCGM ๋ฉํธ๋ฆญ ๋ถ์ฌ
References