| name | coreweave-deploy-integration |
| description | Deploy inference services on CoreWeave with Helm charts and Kustomize.
Use when deploying multi-model inference, managing GPU deployments at scale,
or templating CoreWeave manifests.
Trigger with phrases like "deploy coreweave", "coreweave helm",
"coreweave kustomize", "coreweave deployment patterns".
|
| allowed-tools | Read, Write, Edit, Bash(helm:*), Bash(kubectl:*), Bash(kustomize:*) |
| version | 1.11.0 |
| license | MIT |
| author | Jeremy Longshore <jeremy@intentsolutions.io> |
| tags | ["saas","gpu-cloud","kubernetes","inference","coreweave"] |
| compatibility | Designed for Claude Code |
CoreWeave Deploy Integration
Community-contributed. Not affiliated with, endorsed by, or sponsored by CoreWeave, Inc. CoreWeave is a registered trademark of CoreWeave, Inc.
Overview
Deploy GPU-accelerated inference services on CoreWeave Kubernetes (CKS). This skill covers containerizing inference workloads with NVIDIA CUDA base images, configuring GPU resource limits and node affinity for A100/H100 scheduling, setting up health checks that validate GPU availability and model loading, and executing rolling updates that respect GPU node draining. CoreWeave's scheduler requires explicit GPU resource requests to place pods on the correct hardware tier.
Docker Configuration
FROM nvidia/cuda:12.4.0-runtime-ubuntu22.04 AS base
RUN apt-get update && apt-get install -y --no-install-recommends \
python3 python3-pip curl && rm -rf /var/lib/apt/lists/*
WORKDIR /app
COPY requirements.txt ./
RUN pip3 install --no-cache-dir -r requirements.txt
FROM base
RUN groupadd -r app && useradd -r -g app app
COPY --chown=app:app src/ ./src/
COPY --chown=app:app models/ ./models/
USER app
EXPOSE 8080
HEALTHCHECK --interval=30s --timeout=10s --retries=3 \
CMD curl -f http://localhost:8080/health || exit 1
CMD ["python3", "src/server.py"]
Environment Variables
export COREWEAVE_API_KEY="cw_xxxxxxxxxxxx"
export COREWEAVE_NAMESPACE="tenant-my-org"
export MODEL_NAME="meta-llama/Llama-3.1-8B-Instruct"
export GPU_TYPE="A100_PCIE_80GB"
export GPU_COUNT="1"
export LOG_LEVEL="info"
export PORT="8080"
Health Check Endpoint
import express from 'express';
import { execSync } from 'child_process';
const app = express();
app.get('/health', async (req, res) => {
{
gpuInfo = ().().();
modelLoaded = globalThis. === ;
(!modelLoaded) ();
res.({ : , : gpuInfo, : process.., : ().() });
} (error) {
res.().({ : , : (error ). });
}
});