Skip to main content

agent-reliability-engineering

Implements fault-tolerance mechanisms for AI agent systems including circuit breakers, exponential backoff retries, graceful degradation, health checks, dead letter queues, and timeout management with observability hooks.

설치로 이동

소스 정보

저장소
paulpas/agent-skill-router
최근 소스 활동
2026년 6월 4일 23:31
감지된 SKILL.md 언어
영어
스타
4
포크
1

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.

SKILL.md 표시 중

SKILL.md
소스 지침 · 읽기 전용 미리보기
name
agent-reliability-engineering
description
Implements fault-tolerance mechanisms for AI agent systems including circuit breakers, exponential backoff retries, graceful degradation, health checks, dead letter queues, and timeout management with observability hooks.
license
MIT
compatibility
opencode
metadata
{"version":"1.0.0","domain":"agent","triggers":"fault tolerance, circuit breaker, retry strategy, exponential backoff, graceful degradation, health check, dead letter queue, timeout management graceful degradation","archetypes":["tactical"],"anti_triggers":["brainstorming","vague ideation","single-agent monolith"],"response_profile":{"verbosity":"low","directive_strength":"high","abstraction_level":"operational"},"role":"implementation","scope":"implementation","output-format":"code","content-types":["code","guidance","config","examples","do-dont"],"related-skills":"agent-architecture-patterns, workflow-patterns, failure-mode-analysis"}
# Agent Reliability Engineering Implements fault-tolerance mechanisms for AI agent systems to ensure graceful operation under partial failure. This skill guides the model in applying circuit breakers, retry strategies, degradation patterns, health monitoring, and observability primitives that keep agent systems operational when external dependencies fail. Reliability is not a single pattern — it is a layered defense spanning detection (health checks), prevention (circuit breakers), recovery (retries with backoff), mitigation (graceful degradation), learning (dead letter queues), and visibility (observability hooks) that together prevent cascading failures from taking down the entire agent system. ## TL;DR Checklist - [ ] Configure circuit breakers per external dependency with domain-appropriate thresholds - [ ] Implement exponential backoff with jitter for all retryable operations - [ ] Design graceful degradation paths for each critical service dependency - [ ] Register health checks (startup, readiness, liveness) in correct lifecycle order - [ ] Set up dead letter queues for failed operations requiring manual review - [ ] Wire observability hooks at every failure boundary for metrics collection --- ## When to Use Use this skill when: - Implementing fault tolerance for an agent system that depends on external APIs or services - Designing retry strategies with appropriate backoff policies for transient failures - Building circuit breakers to prevent cascading failures across agent dependencies - Creating graceful degradation paths when downstream services are unavailable - Setting up health checks and monitoring for agent lifecycle management - Implementing dead letter queues for failed operations that need later analysis or reprocessing - Adding observability hooks (metrics, tracing, logging) at failure boundaries ## When NOT to Use Avoid this skill for: - Internal in-process failures — use exception handling patterns instead of circuit breakers - Operations where idempotency guarantees exist and simple retry without backoff suffices - Read-only operations with cached results — add cache invalidation strategies instead - As a substitute for fixing root causes — reliability patterns mask symptoms, they do not cure them --- ## Core Workflow 1. **Classify External Dependencies** — Categorize every external dependency by failure impact and recovery expectation: - Critical: System cannot function without this (e.g., primary LLM provider). Requires circuit breaker + retry + degradation path. - Important: System degrades but remains operational (e.g., secondary data source). Requires circuit breaker + retry with shorter timeout. - Nice-to-have: System continues fully without this (e.g., telemetry enrichment). Requires retry with aggressive timeout only. 2. **Configure Circuit Breakers** — Set per-dependency thresholds based on dependency classification: - Critical dependencies: `failure_threshold=3`, `recovery_timeout=30s`, `half_open_max_calls=1` - Important dependencies: `failure_threshold=5`, `recovery_timeout=60s`, `half_open_max_calls=3` - Nice-to-have dependencies: `failure_threshold=10`, `recovery_timeout=120s`, `half_open_max_calls=5` 3. **Implement Retry with Backoff** — Apply the appropriate retry policy per dependency type: - Transient errors (5xx, timeouts): exponential backoff with jitter starting at 1 second, max 5 retries, max 60 seconds total - Idempotent operations: fixed-window retry up to 3 attempts with 2-second intervals - Non-idempotent operations (writes, trades): single retry with 1-second delay only — never auto-retry without explicit confirmation 4. **Design Degradation Paths** — For each critical dependency, define what "graceful degradation" means: - Primary LLM down → fall back to cached responses or simplified model - Market data feed unavailable → use last-known prices with staleness markers - External API rate-limited → queue requests and retry at reduced throughput 5. **Register Health Checks** — Implement three-level health check hierarchy: - Startup: verify local resources (file access, port binding, config validity) - Readiness: verify all critical dependencies are reachable and responding - Liveness: periodic heartbeat that confirms the process is not in a zombie state 6. **Wire Observability Hooks** — At every failure boundary, emit structured telemetry: - Failure events with dependency name, error code, attempt count, retry delay - Circuit breaker state transitions (closed → open → half_open → closed) - Health check results at each level (startup/ready/live) with timestamps - Dead letter queue depths and processing rates --- ## Implementation Patterns ### Pattern 1: Circuit Breaker with State Machine ```python from enum import Enum import time import threading class CircuitState(Enum): CLOSED = "closed" # Normal operation — requests flow freely OPEN = "open" # Failure threshold exceeded — requests fail fast HALF_OPEN = "half_open" # Recovery probe — limited requests allowed to test restoration class CircuitBreakerError(Exception): """Raised when circuit breaker is open and request is rejected.""" def __init__(self, dependency: str, state: CircuitState) -> None: self.dependency = dependency self.state = state super().__init__(f"Circuit breaker for '{dependency}' is {state.value}") class CircuitBreaker: """State-machine circuit breaker with configurable thresholds. Transitions: CLOSED → OPEN: when failure_count >= failure_threshold within monitoring_window OPEN → HALF_OPEN: when recovery_timeout_seconds has elapsed since opening HALF_OPEN → CLOSED: when half_open_successes successes occur in a row HALF_OPEN → OPEN: on any failure during half-open probing """ def __init__( self, dependency: str, failure_threshold: int = 5, recovery_timeout_seconds: float = 30.0, half_open_max_calls: int = 1, half_open_successes: int = 2, monitoring_window_seconds: float = 60.0, ) -> None: self.dependency = dependency self.failure_threshold = failure_threshold self.recovery_timeout_seconds = recovery_timeout_seconds self.half_open_max_calls = half_open_max_calls self.half_open_successes = half_open_successes self._state = CircuitState.CLOSED self._failure_count = 0 self._success_count_since_failure = 0 self._half_open_calls = 0 self._opened_at: float = 0.0 self._lock = threading.Lock() self._history: list[dict] = [] @property def state(self) -> CircuitState: """Current circuit breaker state with auto-transition from OPEN to HALF_OPEN.""" with self._lock: if self._state == CircuitState.OPEN: elapsed = time.monotonic() - self._opened_at if elapsed >= self.recovery_timeout_seconds: self._transition(CircuitState.HALF_OPEN) return self._state def _transition(self, new_state: CircuitState) -> None: """Record state transition in history and reset per-state counters.""" old_state = self._state self._state = new_state self._history.append({ "from": old_state.value, "to": new_state.value, "timestamp": time.time(), }) if new_state == CircuitState.CLOSED: self._failure_count = 0 self._success_count_since_failure = 0 self._half_open_calls = 0 elif new_state == CircuitState.OPEN: self._opened_at = time.monotonic() self._half_open_calls = 0 elif new_state == CircuitState.HALF_OPEN: self._half_open_calls = 0 def record_success(self) -> None: """Record a successful request. May transition HALF_OPEN → CLOSED.""" with self._lock: if self._state == CircuitState.HALF_OPEN: self._success_count_since_failure += 1 if self._success_count_since_failure >= self.half_open_successes: logger.info( "CircuitBreaker(%s): %d successes in half-open → CLOSED", self.dependency, self.half_open_successes, ) self._transition(CircuitState.CLOSED) def record_failure(self) -> None: """Record a failed request. May transition to OPEN or stay OPEN.""" with self._lock: if self._state == CircuitState.HALF_OPEN: logger.warning( "CircuitBreaker(%s): failure in half-open → OPEN", self.dependency, ) self._transition(CircuitState.OPEN) return self._failure_count += 1 if self._failure_count >= self.failure_threshold: logger.warning( "CircuitBreaker(%s): %d failures in %.0fs window → OPEN", self.dependency, self.failure_count, self.monitoring_window_seconds, ) self._transition(CircuitState.OPEN) def allow_request(self) -> bool: """Check if a request should be allowed through the circuit breaker. Returns True if the request may proceed. Raises CircuitBreakerError if rejected. """ current_state = self.state # Triggers auto-transition check if current_state == CircuitState.CLOSED: return True if current_state == CircuitState.HALF_OPEN: with self._lock: if self._half_open_calls < self.half_open_max_calls: self._half_open_calls += 1 return True raise CircuitBreakerError(self.dependency, CircuitState.HALF_OPEN) # OPEN state — reject immediately raise CircuitBreakerError(self.dependency, CircuitState.OPEN) def get_status(self) -> dict: """Return current circuit breaker status as a serializable dict.""" return { "dependency": self.dependency, "state": self.state.value, "failure_count": self._failure_count, "success_count_since_failure": self._success_count_since_failure, "half_open_calls": self._half_open_calls, "transition_history": self._history[-10:], # Last 10 transitions } def reset(self) -> None: """Manually reset circuit breaker to closed state.""" with self._lock: self._transition(CircuitState.CLOSED) # Example usage: per-dependency circuit breakers in an agent system llm_circuit = CircuitBreaker( dependency="openai_api", failure_threshold=3, recovery_timeout_seconds=30.0, half_open_max_calls=1, half_open_successes=2, ) data_circuit = CircuitBreaker( dependency="market_data_feed", failure_threshold=5, recovery_timeout_seconds=60.0, half_open_max_calls=3, half_open_successes=2, ) async def call_llm_with_breaker(prompt: str) -> str: """LLM call wrapped with circuit breaker protection.""" if not llm_circuit.allow_request(): raise CircuitBreakerError("openai_api", llm_circuit.state) try: response = await fetch_llm_response(prompt) # external API call llm_circuit.record_success() return response except Exception as e: llm_circuit.record_failure() raise async def fetch_market_data(symbol: str) -> dict: """Market data fetch with its own circuit breaker.""" if not data_circuit.allow_request(): # Degradation path: return last-known price instead of failing return {"symbol": symbol, "price": None, "source": "cached_stale"} try: data = await fetch_from_feed(symbol) # external API call data_circuit.record_success() return data except Exception: data_circuit.record_failure() raise ``` ### Pattern 2: Exponential Backoff Retry with Jitter ```python import asyncio import random from typing import Any, Callable, TypeVar T = TypeVar("T") class MaxRetriesExceededError(Exception): """Raised when all retry attempts have been exhausted.""" def __init__( self, operation: str, last_error: Exception, total_attempts: int, total_elapsed_seconds: float, ) -> None: self.operation = operation self.last_error = last_error self.total_attempts = total_attempts self.total_elapsed_seconds = total_elapsed_seconds super().__init__( f"Operation '{operation}' failed after {total_attempts} attempts " f"in {total_elapsed_seconds:.1f}s. Last error: {last_error}" ) class RetryPolicy: """Configurable retry policy with exponential backoff and jitter. Backoff formula: min(max_delay, base_delay * (exponent_factor ^ attempt) + random_jitter) The jitter prevents thundering herd problems when many agents retry simultaneously. Random jitter is uniformly distributed between 0 and base_delay * 0.5. """ def __init__( self, max_retries: int = 3, base_delay: float = 1.0, max_delay: float = 60.0, exponent_factor: float = 2.0, jitter_base: float | None = None, retryable_exceptions: tuple[type[Exception], ...] | None = None, ) -> None: self.max_retries = max_retries self.base_delay = base_delay self.max_delay = max_delay self.exponent_factor = exponent_factor self.jitter_base = jitter_base if jitter_base is not None else base_delay * 0.5 self.retryable_exceptions = retryable_exceptions or ( ConnectionError, TimeoutError, OSError, RuntimeError ) def calculate_delay(self, attempt: int) -> float: """Calculate delay for a given attempt number (0-indexed).""" exponential = self.base_delay * (self.exponent_factor ** attempt) jitter = random.uniform(0, self.jitter_base) return min(self.max_delay, exponential + jitter) async def execute( self, operation: str, fn: Callable[..., Any], *args: Any, **kwargs: Any ) -> T: """Execute an async function with retry logic. Args: operation: Human-readable name for error reporting. fn: Async callable to execute with retries. *args, **kwargs: Arguments passed to the callable. Returns: The result of fn(*args, **kwargs) on success. Raises: MaxRetriesExceededError: After all retry attempts are exhausted. """ last_error: Exception | None = None total_start = asyncio.get_event_loop().time() for attempt in range(self.max_retries + 1): try: if asyncio.iscoroutinefunction(fn): return await fn(*args, **kwargs) else: return fn(*args, **kwargs) # type: ignore[misc] except self.retryable_exceptions as e: last_error = e if attempt < self.max_retries: delay = self.calculate_delay(attempt) logger.info( "RetryPolicy(%s): attempt %d failed (%s), retrying in %.1fs", operation, attempt + 1, type(e).__name__, delay, ) await asyncio.sleep(delay) else: elapsed = asyncio.get_event_loop().time() - total_start raise MaxRetriesExceededError( operation, e, self.max_retries + 1, elapsed, ) from e except Exception as e: # Non-retryable exception — fail immediately raise class FixedWindowRetryPolicy(RetryPolicy): """Fixed-window retry policy for idempotent operations. Unlike exponential backoff, uses a constant delay between retries. Appropriate when the operation is idempotent and rapid recovery is desired. """ def __init__(self, max_retries: int = 3, fixed_delay: float = 2.0) -> None: super().__init__( max_retries=max_retries, base_delay=fixed_delay, max_delay=fixed_delay, exponent_factor=1.0, # No exponential growth jitter_base=fixed_delay * 0.1, # Small jitter only ) # Example usage: retry strategies for different operation types async def fetch_with_retry(url: str) -> dict: """Fetch remote data with exponential backoff — for transient failures.""" policy = RetryPolicy( max_retries=5, base_delay=1.0, max_delay=30.0, exponent_factor=2.0, jitter_base=0.5, ) return await policy.execute(
GitHub에서 보기
이 SKILL.md는 매우 커서 SkillsMP가 여기에는 첫 섹션만 미리 보여줍니다. GitHub에서 보기