Skip to main content

production-error-handling

Production error handling patterns — error taxonomy, retry with exponential backoff, circuit breakers, graceful degradation, dead-letter queues, and structured error logging. Use this skill when the user writes error-prone code (external API calls, database operations, file I/O, network calls), implements retry logic, or asks about resilience patterns. Also trigger when user says /production errors.

Zur Installation springen

Quellinformationen

Repository
vstorm-co/production-stack-skills
Letzte Quellaktivität
16. April 2026 um 17:54
Erkannte Sprache von SKILL.md
Englisch
Sterne
25
Forks
7

Installationsoptionen

Standardmäßig ist der Prompt ausgewählt, der zuerst die Quelle prüft. Sie können zu einem direkten Befehl wechseln oder eine lokale Kopie herunterladen.

Quelldateien prüfen

Lesen Sie SKILL.md und alle von SkillsMP angezeigten Begleitdateien, bevor Sie sich für eine Installation entscheiden.

Datei-Explorer
2 Dateien

SKILL.md wird angezeigt

SKILL.md
Quellanweisungen · Schreibgeschützte Vorschau
name
production-error-handling
description
Production error handling patterns — error taxonomy, retry with exponential backoff, circuit breakers, graceful degradation, dead-letter queues, and structured error logging. Use this skill when the user writes error-prone code (external API calls, database operations, file I/O, network calls), implements retry logic, or asks about resilience patterns. Also trigger when user says /production errors.
# Production Error Handling This skill encodes battle-tested error handling patterns for systems that must stay up when everything around them is falling apart. Every pattern here comes from real production incidents: the retry storm that turned a blip into a 4-hour outage, the bare `except: pass` that silently ate data for three weeks, the missing timeout that let a dead service hold open 200 connections until the pool starved. Follow this guide and none of that happens on your watch. --- ## 1. Error Taxonomy **Classify every error BEFORE writing handling code.** Different errors demand different responses. Treating them the same is how you turn a recoverable hiccup into a cascading outage. ### The Four Categories | Category | Response | Retry? | Alert? | Examples | |----------|----------|--------|--------|----------| | **Transient** | Retry with backoff | Yes | After N failures | Network timeout, 503, connection reset, rate limited (429) | | **Permanent** | Fail immediately | Never | On unexpected frequency | 400, 401, 404, validation error, malformed input | | **Partial** | Degrade gracefully | Optional | Low priority | Cache miss, analytics down, email service down | | **Fatal** | Crash fast | Never | Immediate (PagerDuty) | Missing config, corrupt state, OOM, disk full | ### Exception Hierarchy for a Service Define your exception hierarchy up front. This is not optional -- without it, every developer invents their own error handling and nothing is consistent. ```python class ServiceError(Exception): """Base for all service errors. Every custom exception inherits from this.""" def __init__(self, message: str, *, error_code: str = "INTERNAL_ERROR"): self.message = message self.error_code = error_code super().__init__(message) # --- Transient: safe to retry --- class TransientError(ServiceError): """Temporary failure. Retry with backoff.""" class NetworkTimeoutError(TransientError): def __init__(self, service: str, timeout_seconds: float): super().__init__( f"{service} timed out after {timeout_seconds}s", error_code="NETWORK_TIMEOUT", ) class ServiceUnavailableError(TransientError): def __init__(self, service: str): super().__init__(f"{service} returned 503", error_code="SERVICE_UNAVAILABLE") class RateLimitedError(TransientError): def __init__(self, service: str, retry_after: int | None = None): self.retry_after = retry_after super().__init__(f"{service} rate limited", error_code="RATE_LIMITED") # --- Permanent: do NOT retry --- class PermanentError(ServiceError): """Unrecoverable failure. Do not retry.""" class ValidationError(PermanentError): def __init__(self, field: str, reason: str): super().__init__( f"Validation failed on {field}: {reason}", error_code="VALIDATION_ERROR", ) class NotFoundError(PermanentError): def __init__(self, resource: str, identifier: str): super().__init__( f"{resource} {identifier} not found", error_code="NOT_FOUND", ) class AuthenticationError(PermanentError): def __init__(self): super().__init__("Authentication failed", error_code="AUTH_FAILED") # --- Partial: degrade gracefully --- class PartialError(ServiceError): """Non-critical failure. Continue with degraded functionality.""" class CacheMissError(PartialError): def __init__(self, key: str): super().__init__(f"Cache miss for {key}", error_code="CACHE_MISS") class NonCriticalServiceError(PartialError): def __init__(self, service: str, reason: str): super().__init__( f"Non-critical service {service} failed: {reason}", error_code="NON_CRITICAL_FAILURE", ) # --- Fatal: crash immediately --- class FatalError(ServiceError): """Unrecoverable system failure. Crash the process.""" class MissingConfigError(FatalError): def __init__(self, config_key: str): super().__init__( f"Required config missing: {config_key}", error_code="MISSING_CONFIG", ) class CorruptStateError(FatalError): def __init__(self, detail: str): super().__init__(f"Corrupt state detected: {detail}", error_code="CORRUPT_STATE") ``` **Rules:** - Every exception carries a machine-readable `error_code` for structured logging and API responses - `TransientError` is the ONLY base class that retry logic should catch - Catching `ServiceError` in a handler tells you "something from our domain went wrong" without mixing in random library exceptions - Fatal errors crash the process. Do not try to recover from corrupt state -- you will make it worse --- ## 2. Retry with Exponential Backoff and Jitter ### Why Linear Retry Kills You Linear retry (sleep 1s, try again, sleep 1s, try again) causes **thundering herd**: when a service recovers, all waiting clients slam it simultaneously at the exact same interval. The service goes down again. Repeat until someone pages the on-call. **The correct formula:** ``` delay = min(base * 2^attempt + random(0, jitter), max_delay) ``` - `base`: starting delay (0.5s-1s) - `attempt`: 0-indexed retry count - `jitter`: random component (0 to base) that desynchronizes clients - `max_delay`: cap to prevent absurd waits (30s-60s) ### Python: tenacity tenacity is the standard Python retry library. It handles backoff, jitter, and conditional retry in a composable way. ```python from tenacity import ( retry, stop_after_attempt, wait_exponential_jitter, retry_if_exception_type, before_sleep_log, after_log, ) import structlog logger = structlog.get_logger() # Basic: retry transient errors with exponential backoff + jitter @retry( retry=retry_if_exception_type(TransientError), wait=wait_exponential_jitter(initial=0.5, max=30, jitter=2), stop=stop_after_attempt(4), # 1 initial + 3 retries before_sleep=before_sleep_log(logger, structlog.stdlib.INFO), ) async def call_payment_service(payment_id: str) -> dict: async with httpx.AsyncClient(timeout=5.0) as client: response = await client.post( f"https://payments.internal/charge/{payment_id}" ) if response.status_code == 429: raise RateLimitedError("payment-service") if response.status_code == 503: raise ServiceUnavailableError("payment-service") if response.status_code >= 400: raise PermanentError(f"Payment API returned {response.status_code}") return response.json() # Advanced: custom retry condition with context logging @retry( retry=retry_if_exception_type((TransientError, ConnectionError, TimeoutError)), wait=wait_exponential_jitter(initial=1, max=60, jitter=5), stop=stop_after_attempt(5), before_sleep=before_sleep_log(logger, structlog.stdlib.WARNING), reraise=True, # Re-raise the last exception if all retries fail ) async def fetch_user_profile(user_id: str) -> dict: """Fetch user profile with full retry protection.""" try: async with httpx.AsyncClient(timeout=httpx.Timeout(10.0, connect=3.0)) as client: response = await client.get(f"https://users.internal/profiles/{user_id}") except httpx.ConnectTimeout: raise NetworkTimeoutError("user-service", timeout_seconds=3.0) except httpx.ReadTimeout: raise NetworkTimeoutError("user-service", timeout_seconds=10.0) except httpx.ConnectError: raise TransientError("user-service connection failed", error_code="CONNECT_FAILED") if response.status_code == 404: raise NotFoundError("user", user_id) # Permanent -- do not retry if response.status_code >= 500: raise ServiceUnavailableError("user-service") # Transient -- retry response.raise_for_status() return response.json() ``` ### Node.js: p-retry ```typescript import pRetry, { AbortError } from "p-retry"; async function callPaymentService(paymentId: string): Promise<PaymentResult> { return pRetry( async () => { const response = await fetch( `https://payments.internal/charge/${paymentId}`, { signal: AbortSignal.timeout(5000) } ); // Permanent errors: abort immediately, do not retry if (response.status === 400 || response.status === 401 || response.status === 404) { throw new AbortError(`Permanent failure: ${response.status}`); } // Transient errors: throw normally, p-retry will retry if (response.status === 429 || response.status >= 500) { throw new Error(`Transient failure: ${response.status}`); } return response.json(); }, { retries: 4, minTimeout: 500, // First retry after ~500ms maxTimeout: 30000, // Cap at 30s factor: 2, // Exponential: 500ms, 1s, 2s, 4s randomize: true, // Adds jitter onFailedAttempt: (error) => { console.warn( `Payment call attempt ${error.attemptNumber} failed. ` + `${error.retriesLeft} retries left.`, ); }, }, ); } ``` ### Retry Rules -- Non-Negotiable - **NEVER retry non-idempotent requests** without an idempotency key. Retrying a `POST /charge` without one can double-charge a customer. - **NEVER retry 4xx errors** (except 429). A 400 Bad Request will still be 400 on the next attempt. - **Always set `max_retries`** (3-5 is typical). Without a cap, a persistent failure retries forever. - **Always set `max_delay`** (30s-60s). Without a cap, exponential backoff reaches absurd delays. - **Always set a timeout on the underlying call.** Retry logic without a timeout just stacks up hanging connections. --- ## 3. Circuit Breaker Pattern A circuit breaker stops calling a failing service. Without it, every request to your service blocks for the timeout duration waiting for a dead dependency. ### State Machine ``` success failure_threshold reached +-----------+ +----------------------+ | | | | v OK | v FAILING | CLOSED -------+-----> OPEN ---------> HALF-OPEN ^ | | | | recovery_timeout | | expires | | +-----> probe ----+ | | +---------- success ------------+ (close circuit) ``` - **Closed**: Normal operation. Failures are counted. - **Open**: Circuit is tripped. All calls fail immediately with a fallback. No requests reach the downstream service. - **Half-Open**: After `recovery_timeout`, one probe request is allowed through. If it succeeds, circuit closes. If it fails, circuit re-opens. ### Python: pybreaker ```python import pybreaker import structlog logger = structlog.get_logger() class CircuitBreakerListener(pybreaker.CircuitBreakerListener): def state_change(self, cb, old_state, new_state): logger.warning( "circuit_breaker_state_change", breaker=cb.name, old_state=old_state.name, new_state=new_state.name, ) def failure(self, cb, exc): logger.warning("circuit_breaker_failure", breaker=cb.name, error=str(exc)) payment_breaker = pybreaker.CircuitBreaker( name="payment-service", fail_max=5, # Open after 5 failures reset_timeout=30, # Try again after 30 seconds exclude=[PermanentError], # Don't count permanent errors as failures listeners=[CircuitBreakerListener()], ) @payment_breaker async def call_payment_service(payment_id: str) -> dict: """Wrapped by circuit breaker. Raises CircuitBreakerError when open.""" async with httpx.AsyncClient(timeout=5.0) as client: response = await client.post(f"https://payments.internal/charge/{payment_id}") if response.status_code >= 500: raise ServiceUnavailableError("payment-service") if response.status_code == 404: raise NotFoundError("payment", payment_id) return response.json() # Calling code with fallback async def process_payment(payment_id: str) -> dict: try: return await call_payment_service(payment_id) except pybreaker.CircuitBreakerError: logger.error("circuit_open", service="payment-service", payment_id=payment_id) # Fallback: queue for later processing await enqueue_for_retry("payment", payment_id) return {"status": "queued", "message": "Payment will be processed shortly"} ``` ### Node.js: opossum
Auf GitHub ansehen
Diese SKILL.md ist sehr gross, daher zeigt SkillsMP hier nur den ersten Abschnitt. Auf GitHub ansehen