- name
- ai-deployment-patterns
- description
- Guides expert-level ai deployment patterns implementation: ai-ml and devops decision frameworks, production-ready patterns, and concrete templates for ai deployment patterns workflows.
Use when the user asks about ai deployment patterns, ai deployment patterns configuration, or ai-ml best practices for ai projects.
Do NOT use when the user needs a different ai ml engineering capability -- check sibling skills in the ai ml engineering subcategory.
- license
- Apache-2.0
- metadata
- {"author":"foundry-skills","version":"1.0.0","tags":"ai-ml devops cloud","category":"ai-machine-learning","subcategory":"ai-ml-engineering","depends":"","disclaimer":"none","difficulty":"intermediate"}
# AI Deployment Patterns
## When to Use
**Use this skill when:**
- User is deploying an ML model or LLM-based system to production and needs to choose a serving architecture (batch inference, real-time API, streaming, edge)
- User needs to decide between blue/green deployments, canary releases, shadow mode testing, or A/B model experiments for rolling out a new model version
- User is designing a multi-model pipeline (ensemble, cascade, router, or fallback chain) and needs guidance on orchestration patterns
- User needs to implement model versioning, artifact management, and rollback procedures for a production ML system
- User is experiencing latency, throughput, or cost problems with a deployed model and needs to apply optimization patterns (batching, caching, quantization, async processing)
- User is building a feature store, model registry, or serving layer and needs to understand integration patterns
- User needs to design the observability stack for a live ML system including drift detection, prediction logging, and alerting thresholds
**Do NOT use this skill when:**
- User needs help with model training, hyperparameter tuning, or experiment tracking -- use the ML Training Workflows skill
- User is designing the data pipeline feeding a model -- use the ML Data Engineering skill
- User needs prompt engineering for an LLM application -- use the Prompt Engineering skill
- User is asking about general Kubernetes or Docker deployment without ML-specific concerns -- use the Container Orchestration skill
- User needs cost optimization for cloud infrastructure broadly -- use the Cloud Cost Optimization skill
- User is asking about MLOps platform selection (Vertex AI vs SageMaker vs Azure ML) without a deployment-specific question -- use the MLOps Platform Evaluation skill
- User needs help with model evaluation metrics or offline benchmarking -- use the Model Evaluation skill
---
## Process
### Step 1: Classify the Deployment Scenario
Before recommending any pattern, establish the concrete constraints.
- **Inference mode:** Is this real-time (latency SLA < 500ms), near-real-time (seconds to minutes), or batch (hours-scale jobs)? These require fundamentally different architectures and cannot be swapped cheaply after launch.
- **Request volume:** Estimate requests per second (RPS). Under 10 RPS favors a simple API server; 10-1000 RPS needs autoscaling; over 1000 RPS requires dedicated inference infrastructure with load balancing and horizontal scaling.
- **Model size class:** Categorize as small (< 500MB, e.g., scikit-learn, small ONNX), medium (500MB--5GB, e.g., BERT-family, ResNet), or large (> 5GB, e.g., LLMs requiring GPU or multi-GPU). Size class determines hosting options and cold-start behavior.
- **Latency SLA:** Determine P50, P95, and P99 targets. A P99 of 200ms is very different from P99 of 2 seconds. Most users conflate average with tail latency -- push for explicit P99 requirements.
- **Stateful vs stateless:** Does the model need conversation history, session context, or per-user state? Stateful serving requires sticky routing or external session storage and eliminates many simple horizontal scaling patterns.
- **Data sensitivity:** PII, PHI, or MNPI in the input/output changes which cloud services are usable, logging policies, and whether a model API call can leave the network boundary.
### Step 2: Select the Core Serving Pattern
Match the classified scenario to the correct serving architecture.
- **REST API serving (synchronous):** Use when latency SLA is under 2 seconds and clients can block on a response. Deploy model behind a REST or gRPC interface. Tools: TorchServe, TensorFlow Serving, Triton Inference Server, BentoML, or a plain FastAPI wrapper for small models. Always expose a `/health` and `/ready` endpoint separately -- liveness and readiness probes have different semantics in Kubernetes.
- **Async/queue-based serving:** Use when requests can tolerate > 2 seconds end-to-end, when inputs arrive in bursts, or when the model is expensive (GPU LLM inference). Pattern: producer pushes to a queue (Kafka, SQS, RabbitMQ), worker pool consumes, results written to a results store (Redis, DynamoDB) with a callback or polling endpoint. This decouples client throughput from model throughput.
- **Batch inference:** Use for scheduled scoring jobs (daily churn prediction, weekly recommendation refresh). Pattern: trigger on schedule or data arrival, load model from registry, read input from data warehouse (BigQuery, Redshift, Snowflake), write predictions to output table, track job metadata. Frameworks: Spark MLlib for distributed scoring, Ray for Python-native parallelism, or simple multiprocessing for small-to-medium datasets.
- **Streaming inference:** Use when predictions must be generated on an event stream (fraud detection on transactions, real-time content moderation). Pattern: Kafka Streams or Flink consumer reads events, model is embedded in the processor or called via sidecar, predictions are emitted downstream. Key constraint: model load time must be << stream consumer timeout.
- **Edge/on-device inference:** Use when latency SLA is under 50ms, network is unreliable, or data cannot leave the device. Requires model export to ONNX, TFLite, or CoreML, plus quantization (INT8 or FP16) to fit device memory. Update mechanism must be designed explicitly -- models cannot update like server-side deployments.
### Step 3: Design the Release and Rollout Strategy
Never deploy a new model version directly to 100% of traffic without a staged rollout.
- **Shadow mode (traffic mirroring):** Route 100% of production traffic to the existing model AND mirror it to the new model without serving the new model's predictions to users. Compare outputs offline. Use this pattern first for any significant model change. Requires storing shadow predictions for comparison -- budget for the storage cost. Run shadow mode for at least 24 hours covering a full traffic cycle.
- **Canary deployment:** After shadow mode validation, route a small percentage (typically 1-5%) of live traffic to the new model. Monitor error rates, latency P99, and business metrics. Define explicit rollback criteria before starting the canary (e.g., "roll back if P99 latency increases by > 20% or error rate exceeds 0.5%"). Use weighted routing in your load balancer or service mesh (Istio, AWS ALB weighted target groups).
- **Blue/green deployment:** Maintain two identical serving environments (blue = current production, green = new version). Switch traffic entirely at the load balancer level. Rollback is instant -- flip traffic back to blue. Costs 2x infrastructure during the overlap window. Preferred for model changes that affect the schema of inputs or outputs, where a gradual canary would create version skew problems.
- **A/B model experiment:** Different from a canary -- A/B testing is for measuring business metric impact, not catching errors. Route users consistently (by user ID hash, not randomly per request) to model A or model B, measure downstream conversions or engagement, run a proper statistical test (minimum detectable effect, power analysis) before declaring a winner. Do not run A/B tests for less than the time needed to reach statistical significance, typically 1-2 full weekly business cycles.
- **Feature flags for model routing:** Use a feature flag system (LaunchDarkly, Unleash, or internal) to control model version routing at runtime without redeployment. This enables instant rollback without a new deployment pipeline run.
### Step 4: Implement Model Versioning and the Model Registry
Model artifacts must be versioned and reproducible before any deployment pattern will work reliably.
- **Artifact storage:** Store model artifacts (weights, serialized objects, preprocessing pipelines) in versioned object storage (S3, GCS) with immutable paths (include git commit SHA and training run ID in the path). Never overwrite a model artifact -- always write to a new path.
- **Registry metadata:** For every registered model version, record: training data version or snapshot date, training code git SHA, evaluation metrics on a held-out test set (accuracy, F1, AUC, RMSE as applicable), model size in MB, expected input schema (feature names and dtypes), expected output schema, and the engineer who promoted it to staging/production.
- **Promotion gates:** Enforce a promotion workflow: train --> registered (staging) --> approved (production). Promotion from staging to production must require a passing evaluation suite and at least one human approval. Automate the gate checks; require manual approval only at the final promotion step.
- **Rollback procedure:** A rollback must be a first-class operation that takes under 5 minutes. Document it as a runbook. The rollback should point serving infrastructure at the previous model version artifact path without changing any application code.
### Step 5: Build the Inference Serving Infrastructure
Design the serving layer with production requirements, not prototype assumptions.
- **Containerize the model:** Package the model and all dependencies into a Docker image with a pinned base image (e.g., `python:3.11.4-slim`, not `python:latest`). The inference container must be stateless -- no local file writes during inference. Model weights should be loaded at container startup from object storage, not baked into the image (unless the image size is acceptable and cold start time is not a constraint).
- **Startup and warmup:** After loading weights, run N warmup inference requests (typically 10-50) with representative input shapes before marking the container ready. This fills JIT caches (TorchScript, XLA) and initializes GPU memory. Without warmup, the first real user requests will have 2-10x higher latency.
- **Dynamic batching:** For GPU-served models, enable dynamic batching to amortize the fixed GPU kernel launch overhead. Triton Inference Server supports this natively with configurable `max_queue_delay_microseconds` and `max_batch_size`. A batch size of 8-32 typically gives 3-8x throughput improvement over individual requests. Set `max_queue_delay_microseconds` to no more than 50% of your P99 latency budget.
- **Resource limits:** Set explicit CPU and memory limits on inference containers. For GPU, pin one replica per GPU device or use MIG (Multi-Instance GPU) partitioning for smaller models. Memory leaks in inference servers are common -- set a maximum request count per worker process and recycle workers (gunicorn `--max-requests` or uvicorn equivalent) to bound memory growth.
- **Horizontal autoscaling:** Scale on GPU utilization (target 70-80%) or on request queue depth, not on CPU. Use KEDA (Kubernetes Event-Driven Autoscaling) for queue-depth-based scaling. Set a minimum replica count of at least 2 for any production service to avoid single points of failure and to allow rolling updates without downtime.
- **Timeouts and circuit breakers:** Set aggressive timeouts at every layer: client timeout, load balancer timeout, model server timeout. Use a circuit breaker (Resilience4j, or Envoy circuit breaker via Istio) that opens after N consecutive failures in T seconds. Define the fallback behavior explicitly -- return a cached result, return a default prediction, or return an error with a specific error code that clients handle gracefully.
### Step 6: Implement Observability for Production ML
ML systems have failure modes that standard application monitoring misses entirely.
- **Prediction logging:** Log every inference request: timestamp, input features (or a hash if PII), output prediction, model version, latency, and a unique request ID. Store logs in a structured format (JSON, Parquet) to a queryable store. Prediction logs are the foundation of all other observability.
- **Data drift detection:** Compute feature distribution statistics (mean, standard deviation, percentiles, null rates, categorical frequencies) on a rolling window of prediction logs. Compare to training distribution statistics stored at training time. Alert when the Jensen-Shannon divergence exceeds 0.1 for any individual feature, or when Population Stability Index (PSI) exceeds 0.2 for a critical feature. Tools: Evidently AI, WhyLogs, Nannyml.
- **Concept drift detection:** Monitor prediction distribution over time (mean predicted probability, predicted class distribution, regression output distribution). A shift in prediction distribution is an early signal of concept drift before labels are available. Use Page-Hinkley test or CUSUM for online drift detection.
- **Business metric dashboards:** Connect model predictions to downstream business outcomes where possible (conversions, fraud caught, churn prevented). Build a dashboard tracking these metrics segmented by model version. This is the only way to measure real model value and catch silent failures where the model degrades without technical errors.
- **Alerting thresholds:** Set alerts on: error rate > 1% (P1), P99 latency SLA breach for > 5 minutes (P1), feature null rate increase > 10 percentage points (P2), PSI > 0.25 on any critical feature (P2), prediction distribution shift > 2 standard deviations from 30-day baseline (P2).
- **Trace propagation:** Propagate the request trace ID from the client all the way through to the model inference log. This makes it possible to debug a specific bad prediction by finding it in both the application logs and the model serving logs.
### Step 7: Design Fallback and Graceful Degradation
Every production ML system must have a defined behavior when the model is unavailable or returning low-confidence predictions.
- **Confidence thresholds:** For classification models, define a minimum confidence score below which the model routes to a fallback. Example: if predicted probability max < 0.6, route to a rule-based system or human review queue instead of serving the low-confidence prediction.
- **Fallback hierarchy:** Design a ranked fallback chain: (1) primary model, (2) smaller/faster backup model, (3) rule-based heuristic, (4) cached most-recent prediction, (5) safe default. The farther down the chain, the more degraded the user experience, but the service remains available.
- **Cache layer for idempotent predictions:** For inputs that are repeated frequently (product recommendations for popular items, content moderation for viral content), cache predictions in Redis with a TTL matched to acceptable staleness (typically 1-60 minutes). Cache hit rate of 30-50% is achievable for recommendation systems and eliminates that fraction of model calls entirely.
- **Timeout budget:** Allocate the total request timeout budget across all stages. If the total SLA is 500ms, the model inference budget might be 300ms, leaving 200ms for preprocessing, postprocessing, and network overhead. If the model does not respond within 300ms, execute the fallback path immediately -- do not wait for the full request timeout.
---
## Output Format
Produce the following artifacts when designing or documenting an AI deployment pattern.
### Deployment Pattern Decision Summary
```
Model Deployment Decision Record
=================================
Date: 2024-01-15
Author: <engineer name>
Model: customer-churn-classifier-v3
Use Case: Real-time churn risk scoring at checkout
Serving Pattern: Synchronous REST API (FastAPI + BentoML)
Rationale: P99 latency SLA = 200ms, RPS = 45 peak, model size = 85MB
Release Strategy: Shadow mode (48h) → Canary 5% (24h) → Canary 25% (24h) → Full
Rollback Trigger: P99 > 180ms OR error rate > 0.5% OR prediction_null_rate > 2%
Fallback: Rule-based threshold (recency + frequency score)
Infrastructure:
Replicas (min/max): 2 / 10
Container CPU: 0.5 / 2.0 cores
Container Memory: 1Gi / 2Gi
Autoscale Metric: request queue depth > 50
Monitoring:
Drift Detection: PSI on 6 input features, 1-hour rolling window
Alert P1: error_rate > 1%, P99 > 200ms
Alert P2: PSI > 0.2, prediction_mean shift > 1.5 std
```
### Architecture Diagram (ASCII)
```
Client Request
|
v
[API Gateway / Load Balancer]
|
|-- Feature Flag: canary? --> [Model v3 Serving Cluster]
| |
|-- (else) -----------------> [Model v2 Serving Cluster]
|
+-- Input Validation (schema check)
+-- Feature Preprocessing
+-- [Model Inference Engine]
| |
| [Prediction Cache (Redis, TTL=5min)]
| |
+-- Confidence Check
|
>= 0.6 -------+------- < 0.6
| |
Serve prediction [Fallback: rule-based]
|
[Prediction Logger]
|
[Response to Client]
Async path:
[Prediction Logger] --> [Kafka topic: model-predictions]
|
[Drift Detector (Evidently)]
|
[Metrics Dashboard (Grafana)]
|
[Alert Manager (PagerDuty)]
```
### Serving Configuration Template (BentoML / FastAPI)
```python
# serving/model_service.py
# Synchronous REST inference service with fallback, caching, and observability
import time
import hashlib
import logging
from typing import Optional
from dataclasses import dataclass
import redis
import bentoml
from bentoml.io import NumpyNdarray, JSON
from prometheus_client import Histogram, Counter, Gauge
# -- Prometheus metrics
INFERENCE_LATENCY = Histogram(
"model_inference_latency_seconds",
"Model inference latency",
buckets=[0.01, 0.025, 0.05, 0.1, 0.2, 0.5, 1.0, 2.0],
labelnames=["model_version", "path"], # path: "model" | "cache" | "fallback"
)
PREDICTION_COUNTER = Counter(
"model_predictions_total",
"Total predictions served",
labelnames=["model_version", "path", "confidence_bucket"],
)
FALLBACK_COUNTER = Counter(
"model_fallback_total",
"Total fallback activations",
labelnames=["reason"], # "low_confidence" | "timeout" | "error"
)
logger = logging.getLogger(__name__)
@dataclass
class InferenceConfig:
model_tag: str # e.g. "churn-classifier:v3.2.1"
model_version: str # e.g. "v3.2.1" for labels
confidence_threshold: float # e.g. 0.60
cache_ttl_seconds: int # e.g. 300
inference_timeout_ms: int # e.g. 150
redis_url: str # e.g. "redis://redis-svc:6379/0"
warmup_requests: int # e.g. 20
class ModelService:
def __init__(self, config: InferenceConfig) -> None:
self.config = config
self.runner = bentoml.models.get(config.model_tag).to_runner()
self.cache = redis.from_url(config.redis_url, decode_responses=True)
self._warmup()
def _warmup(self) -> None:
"""Run warmup inferences before marking service ready."""
import numpy as np
dummy_input = np.zeros((1, 12), dtype=np.float32) # match training feature count
for i in range(self.config.warmup_requests):
self.runner.predict.run(dummy_input)
logger.info(
"Warmup complete",
extra={"warmup_requests": self.config.warmup_requests}
)
def _cache_key(self, features: dict) -> str:
canonical = str(sorted(features.items()))
return f"pred:{hashlib.sha256(canonical.encode()).hexdigest()[:16]}"
def _rule_based_fallback(self, features: dict) -> dict:
"""
Rule-based fallback: deterministic churn score from recency + frequency.
Returns same schema as model output for transparent substitution.
"""
recency_days = features.get("days_since_last_order", 999)
order_count = features.get("order_count_90d", 0)
score = min(1.0, recency_days / 180.0) * (1.0 - min(1.0, order_count / 10.0))
return {
"churn_probability": round(score, 4),
"confidence": 0.0, # signal to downstream that this is a fallback
"prediction_source": "fallback_rules",
}
async def predict(self, features: dict) -> dict:
start = time.perf_counter()
cache_key = self._cache_key(features)
# -- Cache check
cached = self.cache.get(cache_key)
if cached:
elapsed = (time.perf_counter() - start) * 1000
INFERENCE_LATENCY.labels(
model_version=self.config.model_version, path="cache"
).observe(elapsed / 1000)
PREDICTION_COUNTER.labels(
model_version=self.config.model_version,
path="cache",
confidence_bucket="high",
).inc()
return {"churn_probability": float(cached), "prediction_source": "cache"}
# -- Model inference with timeout guard
import asyncio
import numpy as np
feature_vector = np.array(
[list(features.values())], dtype=np.float32
)
try:
result = await asyncio.wait_for(
self.runner.predict.async_run(feature_vector),
timeout=self.config.inference_timeout_ms / 1000.0,
)
churn_prob = float(result[0][1]) # index 1 = churn class probability
confidence = max(churn_prob, 1.0 - churn_prob)
elapsed = (time.perf_counter() - start) * 1000
confidence_bucket = "high" if confidence >= 0.6 else "low"
INFERENCE_LATENCY.labels(
model_version=self.config.model_version, path="model"
Auf GitHub ansehen