Skip to main content

production-error-handling

Production error handling patterns — error taxonomy, retry with exponential backoff, circuit breakers, graceful degradation, dead-letter queues, and structured error logging. Use this skill when the user writes error-prone code (external API calls, database operations, file I/O, network calls), implements retry logic, or asks about resilience patterns. Also trigger when user says /production errors.

Aller à l'installation

Informations de source

Dépôt
vstorm-co/production-stack-skills
Dernière activité de la source
16 avril 2026 à 17:54
Langue détectée de SKILL.md
anglais
Étoiles
25
Forks
7

Options d'installation

Le prompt qui vérifie d'abord la source est sélectionné par défaut. Vous pouvez passer à une commande directe ou télécharger une copie locale.

Vérifiez les fichiers source

Lisez SKILL.md et les fichiers associés affichés par SkillsMP avant de décider de l'installer.

Explorateur de fichiers
2 fichiers

Affichage de SKILL.md

SKILL.md
Instructions source · Aperçu en lecture seule
name
production-error-handling
description
Production error handling patterns — error taxonomy, retry with exponential backoff, circuit breakers, graceful degradation, dead-letter queues, and structured error logging. Use this skill when the user writes error-prone code (external API calls, database operations, file I/O, network calls), implements retry logic, or asks about resilience patterns. Also trigger when user says /production errors.
# Production Error Handling This skill encodes battle-tested error handling patterns for systems that must stay up when everything around them is falling apart. Every pattern here comes from real production incidents: the retry storm that turned a blip into a 4-hour outage, the bare `except: pass` that silently ate data for three weeks, the missing timeout that let a dead service hold open 200 connections until the pool starved. Follow this guide and none of that happens on your watch. --- ## 1. Error Taxonomy **Classify every error BEFORE writing handling code.** Different errors demand different responses. Treating them the same is how you turn a recoverable hiccup into a cascading outage. ### The Four Categories | Category | Response | Retry? | Alert? | Examples | |----------|----------|--------|--------|----------| | **Transient** | Retry with backoff | Yes | After N failures | Network timeout, 503, connection reset, rate limited (429) | | **Permanent** | Fail immediately | Never | On unexpected frequency | 400, 401, 404, validation error, malformed input | | **Partial** | Degrade gracefully | Optional | Low priority | Cache miss, analytics down, email service down | | **Fatal** | Crash fast | Never | Immediate (PagerDuty) | Missing config, corrupt state, OOM, disk full | ### Exception Hierarchy for a Service Define your exception hierarchy up front. This is not optional -- without it, every developer invents their own error handling and nothing is consistent. ```python class ServiceError(Exception): """Base for all service errors. Every custom exception inherits from this.""" def __init__(self, message: str, *, error_code: str = "INTERNAL_ERROR"): self.message = message self.error_code = error_code super().__init__(message) # --- Transient: safe to retry --- class TransientError(ServiceError): """Temporary failure. Retry with backoff.""" class NetworkTimeoutError(TransientError): def __init__(self, service: str, timeout_seconds: float): super().__init__( f"{service} timed out after {timeout_seconds}s", error_code="NETWORK_TIMEOUT", ) class ServiceUnavailableError(TransientError): def __init__(self, service: str): super().__init__(f"{service} returned 503", error_code="SERVICE_UNAVAILABLE") class RateLimitedError(TransientError): def __init__(self, service: str, retry_after: int | None = None): self.retry_after = retry_after super().__init__(f"{service} rate limited", error_code="RATE_LIMITED") # --- Permanent: do NOT retry --- class PermanentError(ServiceError): """Unrecoverable failure. Do not retry.""" class ValidationError(PermanentError): def __init__(self, field: str, reason: str): super().__init__( f"Validation failed on {field}: {reason}", error_code="VALIDATION_ERROR", ) class NotFoundError(PermanentError): def __init__(self, resource: str, identifier: str): super().__init__( f"{resource} {identifier} not found", error_code="NOT_FOUND", ) class AuthenticationError(PermanentError): def __init__(self): super().__init__("Authentication failed", error_code="AUTH_FAILED") # --- Partial: degrade gracefully --- class PartialError(ServiceError): """Non-critical failure. Continue with degraded functionality.""" class CacheMissError(PartialError): def __init__(self, key: str): super().__init__(f"Cache miss for {key}", error_code="CACHE_MISS") class NonCriticalServiceError(PartialError): def __init__(self, service: str, reason: str): super().__init__( f"Non-critical service {service} failed: {reason}", error_code="NON_CRITICAL_FAILURE", ) # --- Fatal: crash immediately --- class FatalError(ServiceError): """Unrecoverable system failure. Crash the process.""" class MissingConfigError(FatalError): def __init__(self, config_key: str): super().__init__( f"Required config missing: {config_key}", error_code="MISSING_CONFIG", ) class CorruptStateError(FatalError): def __init__(self, detail: str): super().__init__(f"Corrupt state detected: {detail}", error_code="CORRUPT_STATE") ``` **Rules:** - Every exception carries a machine-readable `error_code` for structured logging and API responses - `TransientError` is the ONLY base class that retry logic should catch - Catching `ServiceError` in a handler tells you "something from our domain went wrong" without mixing in random library exceptions - Fatal errors crash the process. Do not try to recover from corrupt state -- you will make it worse --- ## 2. Retry with Exponential Backoff and Jitter ### Why Linear Retry Kills You Linear retry (sleep 1s, try again, sleep 1s, try again) causes **thundering herd**: when a service recovers, all waiting clients slam it simultaneously at the exact same interval. The service goes down again. Repeat until someone pages the on-call. **The correct formula:** ``` delay = min(base * 2^attempt + random(0, jitter), max_delay) ``` - `base`: starting delay (0.5s-1s) - `attempt`: 0-indexed retry count - `jitter`: random component (0 to base) that desynchronizes clients - `max_delay`: cap to prevent absurd waits (30s-60s) ### Python: tenacity tenacity is the standard Python retry library. It handles backoff, jitter, and conditional retry in a composable way. ```python from tenacity import ( retry, stop_after_attempt, wait_exponential_jitter, retry_if_exception_type, before_sleep_log, after_log, ) import structlog logger = structlog.get_logger() # Basic: retry transient errors with exponential backoff + jitter @retry( retry=retry_if_exception_type(TransientError), wait=wait_exponential_jitter(initial=0.5, max=30, jitter=2), stop=stop_after_attempt(4), # 1 initial + 3 retries before_sleep=before_sleep_log(logger, structlog.stdlib.INFO), ) async def call_payment_service(payment_id: str) -> dict: async with httpx.AsyncClient(timeout=5.0) as client: response = await client.post( f"https://payments.internal/charge/{payment_id}" ) if response.status_code == 429: raise RateLimitedError("payment-service") if response.status_code == 503: raise ServiceUnavailableError("payment-service") if response.status_code >= 400: raise PermanentError(f"Payment API returned {response.status_code}") return response.json() # Advanced: custom retry condition with context logging @retry( retry=retry_if_exception_type((TransientError, ConnectionError, TimeoutError)), wait=wait_exponential_jitter(initial=1, max=60, jitter=5), stop=stop_after_attempt(5), before_sleep=before_sleep_log(logger, structlog.stdlib.WARNING), reraise=True, # Re-raise the last exception if all retries fail ) async def fetch_user_profile(user_id: str) -> dict: """Fetch user profile with full retry protection.""" try: async with httpx.AsyncClient(timeout=httpx.Timeout(10.0, connect=3.0)) as client: response = await client.get(f"https://users.internal/profiles/{user_id}") except httpx.ConnectTimeout: raise NetworkTimeoutError("user-service", timeout_seconds=3.0) except httpx.ReadTimeout: raise NetworkTimeoutError("user-service", timeout_seconds=10.0) except httpx.ConnectError: raise TransientError("user-service connection failed", error_code="CONNECT_FAILED") if response.status_code == 404: raise NotFoundError("user", user_id) # Permanent -- do not retry if response.status_code >= 500: raise ServiceUnavailableError("user-service") # Transient -- retry response.raise_for_status() return response.json() ``` ### Node.js: p-retry ```typescript import pRetry, { AbortError } from "p-retry"; async function callPaymentService(paymentId: string): Promise<PaymentResult> { return pRetry( async () => { const response = await fetch( `https://payments.internal/charge/${paymentId}`, { signal: AbortSignal.timeout(5000) } ); // Permanent errors: abort immediately, do not retry if (response.status === 400 || response.status === 401 || response.status === 404) { throw new AbortError(`Permanent failure: ${response.status}`); } // Transient errors: throw normally, p-retry will retry if (response.status === 429 || response.status >= 500) { throw new Error(`Transient failure: ${response.status}`); } return response.json(); }, { retries: 4, minTimeout: 500, // First retry after ~500ms maxTimeout: 30000, // Cap at 30s factor: 2, // Exponential: 500ms, 1s, 2s, 4s randomize: true, // Adds jitter onFailedAttempt: (error) => { console.warn( `Payment call attempt ${error.attemptNumber} failed. ` + `${error.retriesLeft} retries left.`, ); }, }, ); } ``` ### Retry Rules -- Non-Negotiable - **NEVER retry non-idempotent requests** without an idempotency key. Retrying a `POST /charge` without one can double-charge a customer. - **NEVER retry 4xx errors** (except 429). A 400 Bad Request will still be 400 on the next attempt. - **Always set `max_retries`** (3-5 is typical). Without a cap, a persistent failure retries forever. - **Always set `max_delay`** (30s-60s). Without a cap, exponential backoff reaches absurd delays. - **Always set a timeout on the underlying call.** Retry logic without a timeout just stacks up hanging connections. --- ## 3. Circuit Breaker Pattern A circuit breaker stops calling a failing service. Without it, every request to your service blocks for the timeout duration waiting for a dead dependency. ### State Machine ``` success failure_threshold reached +-----------+ +----------------------+ | | | | v OK | v FAILING | CLOSED -------+-----> OPEN ---------> HALF-OPEN ^ | | | | recovery_timeout | | expires | | +-----> probe ----+ | | +---------- success ------------+ (close circuit) ``` - **Closed**: Normal operation. Failures are counted. - **Open**: Circuit is tripped. All calls fail immediately with a fallback. No requests reach the downstream service. - **Half-Open**: After `recovery_timeout`, one probe request is allowed through. If it succeeds, circuit closes. If it fails, circuit re-opens. ### Python: pybreaker ```python import pybreaker import structlog logger = structlog.get_logger() class CircuitBreakerListener(pybreaker.CircuitBreakerListener): def state_change(self, cb, old_state, new_state): logger.warning( "circuit_breaker_state_change", breaker=cb.name, old_state=old_state.name, new_state=new_state.name, ) def failure(self, cb, exc): logger.warning("circuit_breaker_failure", breaker=cb.name, error=str(exc)) payment_breaker = pybreaker.CircuitBreaker( name="payment-service", fail_max=5, # Open after 5 failures reset_timeout=30, # Try again after 30 seconds exclude=[PermanentError], # Don't count permanent errors as failures listeners=[CircuitBreakerListener()], ) @payment_breaker async def call_payment_service(payment_id: str) -> dict: """Wrapped by circuit breaker. Raises CircuitBreakerError when open.""" async with httpx.AsyncClient(timeout=5.0) as client: response = await client.post(f"https://payments.internal/charge/{payment_id}") if response.status_code >= 500: raise ServiceUnavailableError("payment-service") if response.status_code == 404: raise NotFoundError("payment", payment_id) return response.json() # Calling code with fallback async def process_payment(payment_id: str) -> dict: try: return await call_payment_service(payment_id) except pybreaker.CircuitBreakerError: logger.error("circuit_open", service="payment-service", payment_id=payment_id) # Fallback: queue for later processing await enqueue_for_retry("payment", payment_id) return {"status": "queued", "message": "Payment will be processed shortly"} ``` ### Node.js: opossum
Voir sur GitHub
Ce SKILL.md est tres volumineux, SkillsMP affiche donc ici seulement la premiere section. Voir sur GitHub