| name | python-observability |
| description | Python observability patterns including structured logging, metrics, and distributed tracing. |
| when_to_use | Adding logging, implementing metrics collection, setting up tracing, or debugging production systems. |
Python Observability
Instrument Python applications with structured logs, metrics, and traces. When something breaks in production, you need to answer "what, where, and why" without deploying new code.
Core Concepts
1. Structured Logging
Emit logs as JSON with consistent fields for production environments. Machine-readable logs enable powerful queries and alerts. For local development, consider human-readable formats.
2. The Four Golden Signals
Track latency, traffic, errors, and saturation for every service boundary.
3. Correlation IDs
Thread a unique ID through all logs and spans for a single request, enabling end-to-end tracing.
4. Bounded Cardinality
Keep metric label values bounded. Unbounded labels (like user IDs) explode storage costs.
Quick Start
import structlog
structlog.configure(
processors=[
structlog.processors.TimeStamper(fmt="iso"),
structlog.processors.JSONRenderer(),
],
)
logger = structlog.get_logger()
logger.info("Request processed", user_id="123", duration_ms=45)
Fundamental Patterns
Pattern 1: Structured Logging with Structlog
Configure structlog for JSON output with consistent fields.
import logging
import structlog
def configure_logging(log_level: str = "INFO") -> None:
"""Configure structured logging for the application."""
structlog.configure(
processors=[
structlog.contextvars.merge_contextvars,
structlog.processors.add_log_level,
structlog.processors.TimeStamper(fmt="iso"),
structlog.processors.StackInfoRenderer(),
structlog.processors.format_exc_info,
structlog.processors.JSONRenderer(),
],
wrapper_class=structlog.make_filtering_bound_logger(
getattr(logging, log_level.upper())
),
context_class=dict,
logger_factory=structlog.PrintLoggerFactory(),
cache_logger_on_first_use=True,
)
configure_logging("INFO")
logger = structlog.get_logger()
Pattern 2: Consistent Log Fields
Every log entry should include standard fields for filtering and correlation.
import structlog
from contextvars import ContextVar
correlation_id: ContextVar[str] = ContextVar("correlation_id", default="")
logger = structlog.get_logger()
def process_request(request: Request) -> Response:
"""Process request with structured logging."""
logger.info(
"Request received",
correlation_id=correlation_id.get(),
method=request.method,
path=request.path,
user_id=request.user_id,
)
try:
result = handle_request(request)
logger.info(
"Request completed",
correlation_id=correlation_id.get(),
status_code=200,
duration_ms=elapsed,
)
return result
except Exception as e:
logger.error(
"Request failed",
correlation_id=correlation_id.get(),
error_type=type(e).__name__,
error_message=str(e),
)
raise
Pattern 3: Semantic Log Levels
Use log levels consistently across the application.
| Level | Purpose | Examples |
|---|
DEBUG | Development diagnostics | Variable values, internal state |
INFO | Request lifecycle, operations | Request start/end, job completion |
WARNING | Recoverable anomalies | Retry attempts, fallback used |
ERROR | Failures needing attention | Exceptions, service unavailable |
logger.debug("Cache lookup", key=cache_key, hit=cache_hit)
logger.info("Order created", order_id=order.id, total=order.total)
logger.warning(
"Rate limit approaching",
current_rate=950,
limit=1000,
reset_seconds=30,
)
logger.error(
"Payment processing failed",
order_id=order.id,
error=str(e),
payment_provider="stripe",
)
Never log expected behavior at ERROR. A user entering a wrong password is INFO, not ERROR.
Pattern 4: Correlation ID Propagation
Generate a unique ID at ingress and thread it through all operations.
from contextvars import ContextVar
import uuid
import structlog
correlation_id: ContextVar[str] = ContextVar("correlation_id", default="")
def set_correlation_id(cid: str | None = None) -> str:
"""Set correlation ID for current context."""
cid = cid or str(uuid.uuid4())
correlation_id.set(cid)
structlog.contextvars.bind_contextvars(correlation_id=cid)
return cid
from fastapi import Request
async def correlation_middleware(request: Request, call_next):
"""Middleware to set and propagate correlation ID."""
cid = request.headers.get("X-Correlation-ID") or str(uuid.uuid4())
set_correlation_id(cid)
response = await call_next(request)
response.headers["X-Correlation-ID"] = cid
return response
Propagate to outbound requests:
import httpx
async def call_downstream_service(endpoint: str, data: dict) -> dict:
"""Call downstream service with correlation ID."""
async with httpx.AsyncClient() as client:
response = await client.post(
endpoint,
json=data,
headers={"X-Correlation-ID": correlation_id.get()},
)
return response.json()
Pattern 5: Deferred Log Message Formatting
Never f-string a log message. Pass the template and its values as separate args so formatting is deferred -- it runs only when the record's level is enabled (Logger.isEnabledFor short-circuits before interpolation; an f-string, or %/str.format/+ pre-formatting, pays that cost even for filtered-out records). Two mechanisms:
stdlib_logger.info("user=%s orders=%s", user_id, order_count)
logger.info("orders processed", user_id=user_id, order_count=order_count)
stdlib_logger.info(f"user={user_id} orders={order_count}")
%s args are not the structlog form: this skill's processor chain (Pattern 1) omits PositionalArgumentsFormatter, so logger.info("user=%s", user_id) stores the literal user=%s under positional_args instead of interpolating. Use kwargs with structlog, %s args with stdlib.
Ruff flake8-logging-format (G) enforces this: G004 flags f-strings, G001/G002/G003 flag str.format/%/+ pre-formatting.
For exceptions use logger.exception (= error(..., exc_info=True)) -- it attaches the traceback; keep deferred %s args:
except OSError:
stdlib_logger.exception("save failed, record_id=%s", record.id)
raise
Detailed worked examples and patterns
Detailed sections (starting with ## Advanced Patterns) live in references/details.md. Read that file when the navigation summary above is insufficient.
Best Practices Summary
- Use structured logging - JSON logs with consistent fields
- Propagate correlation IDs - Thread through all requests and logs
- Track the four golden signals - Latency, traffic, errors, saturation
- Bound label cardinality - Never use unbounded values as metric labels
- Log at appropriate levels - Don't cry wolf with ERROR
- Include context - User ID, request ID, operation name in logs
- Use context managers - Consistent timing and error handling
- Separate concerns - Observability code shouldn't pollute business logic
- Test your observability - Verify logs and metrics in integration tests
- Set up alerts - Metrics are useless without alerting
- Defer log formatting - Pass
%s args (stdlib) or kwargs (structlog); never f-string a log message (ruff G004)