- name
- mcp-production-patterns
- description
- Implements production-grade MCP deployments with patterns for latency optimization (stdio vs HTTP transports), stateful session management, caching strategies, observability (tracing, metrics), rate limiting, graceful degradation, and lessons from production evidence (Smartsheet, Rippling, Chronograph deployments).
- license
- MIT
- compatibility
- opencode
- metadata
- {"version":"1.0.0","domain":"cncf","role":"implementation","scope":"infrastructure","output-format":"code","triggers":"mcp production, mcp deployment, mcp performance, mcp observability, mcp monitoring, http streamable, stateful session pattern","related-skills":"mcp-server-fastmcp-python, mcp-client-integration, mcp-security-authorization","archetypes":"tactical, strategic","anti_triggers":"brainstorming, vague ideation","response_profile":{"verbosity":"medium","directive_strength":"high","abstraction_level":"operational"}}
# MCP Production Patterns
Implements production-grade Model Context Protocol deployments with proven patterns for transport optimization, session management, caching, observability, and graceful degradation based on real-world evidence from Smartsheet, Rippling, and Chronograph deployments.
## TL;DR Checklist
- [ ] Choose transport: stdio (0.01ms local) vs HTTP/Streamable (0.39ms loopback, 2.5s coldstart)
- [ ] Apply one of five core server patterns: Resource Gateway, Tool Orchestrator, Stateful Session, Proxy Aggregator, Domain-Specific Adapter
- [ ] Implement caching for tool results and resource listings (avoid cache stampedes)
- [ ] Add distributed tracing (OpenTelemetry) and structured metrics
- [ ] Implement rate limiting (token bucket with adaptive backoff)
- [ ] Add circuit breaker for failing backends
- [ ] Test graceful degradation: timeouts, fallbacks, partial responses
---
## When to Use
Use this skill when:
- Deploying MCP servers to production or high-traffic environments
- Optimizing for latency-sensitive applications (sub-100ms response requirement)
- Building multi-tenant MCP deployments (Rippling pattern)
- Aggregating resources from multiple backend systems (Proxy Aggregator)
- Implementing stateful conversation context across multiple MCP calls
- Adding observability (tracing, metrics, structured logging) to MCP infrastructure
- Designing rate limiting or graceful degradation strategies
- Choosing between stdio, HTTP, or Streamable transports based on deployment topology
---
## When NOT to Use
Avoid this skill for:
- Simple, local MCP servers (use default patterns)
- Single-use development/testing servers
- Deployments with <10 QPS traffic
- Scenarios where you don't need cross-call state management
- When native MCP transport defaults are sufficient (not performance-constrained)
---
## Transport Selection: Latency Profile & Trade-offs
### Latency Characteristics
| Transport | Latency | Cold Start | Best For | Trade-off |
|-----------|---------|-----------|----------|-----------|
| **stdio** | 0.01ms (local process) | Instant | Single-machine, local LLM agents | Limited to same machine; hard to scale |
| **HTTP/REST** | 0.39ms (loopback), ~5-50ms (network) | Instant | Distributed deployments, cloud native | Network overhead; connection pooling critical |
| **Streamable (SSE-based)** | 0.39ms (loopback), 2.5s (cold connection) | 2.5s | Long-lived connections, real-time updates | Higher memory per connection; reconnect storms |
### Selection Logic
```
Is the LLM agent on the same machine?
→ YES: Use stdio (simplest, fastest)
→ NO: Is real-time streaming required?
→ YES: Use Streamable (handle reconnection storms)
→ NO: Use HTTP/REST (standard, cloud-native, easiest load balancing)
```
### HTTP-Specific Tuning
For HTTP deployments, always implement:
1. **Connection pooling**: Reuse TCP connections
2. **Keep-Alive headers**: `Connection: keep-alive`
3. **Graceful shutdown**: Drain in-flight requests before closing
4. **Load balancing**: Distribute across multiple server instances
5. **Circuit breaker**: Fail fast if backend is overloaded
---
## Five Core Server Patterns
### Pattern 1: Resource Gateway
**Purpose:** Expose read-only resources from multiple backend systems. Delegate tool invocation to backend-specific servers.
**Real-world example:** Smartsheet MCP integration — expose sheets, rows, columns as resources; delegate writes to specialized tool server.
**When to use:**
- Resources are expensive to compute but stable (low churn)
- Tools require separate authorization or backend delegation
- You want to cache resource listings aggressively
**Implementation:**
```python
from mcp.server import Server
from mcp.types import Resource, ResourceTemplate, ListResourcesRequest
server = Server("resource-gateway")
# Cache resource listings (30-second TTL)
_resource_cache = {}
_cache_ttl = 30
@server.list_resources()
async def list_resources_handler(request: ListResourcesRequest) -> list[Resource]:
"""Expose resources from multiple backends."""
now = time.time()
# Check cache first
cache_key = f"resources:{request.cursor or 'root'}"
if cache_key in _resource_cache:
cached, cached_at = _resource_cache[cache_key]
if now - cached_at < _cache_ttl:
return cached # Cache hit
# Cache miss: fetch from backends
resources = []
# Backend 1: Smartsheet API
smartsheet_resources = await fetch_smartsheet_resources()
resources.extend(smartsheet_resources)
# Backend 2: Airtable API
airtable_resources = await fetch_airtable_resources()
resources.extend(airtable_resources)
# Cache the result
_resource_cache[cache_key] = (resources, now)
return resources
@server.call_tool()
async def call_tool_handler(name: str, arguments: dict) -> list:
"""Delegate tool execution to specialized servers."""
if name.startswith("smartsheet_"):
# Route to Smartsheet tool server
return await smartsheet_tool_server.call_tool(name, arguments)
elif name.startswith("airtable_"):
# Route to Airtable tool server
return await airtable_tool_server.call_tool(name, arguments)
else:
raise ValueError(f"Unknown tool: {name}")
```
**Trade-offs:**
- ✅ Clean separation: read-only gateway + write-capable tool servers
- ✅ Resource listing is cached and fast
- ❌ Tool invocation adds hop latency
- ❌ Gateway must know about all backends
---
### Pattern 2: Tool Orchestrator
**Purpose:** Receive multi-step tool requests, route them to appropriate backend servers, and return aggregated results.
**Real-world example:** Rippling multi-tenant integration — route HRIS calls to Rippling, payroll calls to ADP, benefit calls to Guidepoint.
**When to use:**
- Tools map to different backend systems
- Tool invocation logic depends on user tenant/organization
- You want a single MCP endpoint that routes to many backends
**Implementation:**
```python
@server.call_tool()
async def call_tool_handler(name: str, arguments: dict) -> list:
"""Multi-tenant tool orchestrator."""
tenant_id = arguments.get("tenant_id")
if not tenant_id:
raise ValueError("tenant_id is required")
# Route based on tool name and tenant
if name == "get_employee":
backend = get_tenant_backend(tenant_id, "hris")
return await backend.call_tool(name, arguments)
elif name == "update_payroll":
backend = get_tenant_backend(tenant_id, "payroll")
# Add circuit breaker for resilience
try:
return await circuit_breaker.call(
backend.call_tool,
name,
arguments,
timeout=5.0
)
except CircuitBreakerOpen:
return [{"error": "Payroll service temporarily unavailable"}]
else:
raise ValueError(f"Unknown tool: {name}")
```
**Trade-offs:**
- ✅ Single endpoint for all tools
- ✅ Easy to add new backends
- ❌ Requires tenant/org context in every call
- ❌ Error in one backend doesn't affect others (good for resilience, bad for atomicity)
---
### Pattern 3: Stateful Session
**Purpose:** Maintain conversation state across multiple tool calls. Store state server-side with session ID.
**Real-world example:** Chronograph time-series database — maintain query context (selected metrics, time range) across multiple tool invocations.
**When to use:**
- Tools need context from previous calls
- You want to reduce redundant data transfer
- Multi-step workflows where state builds up
**Implementation:**
```python
from dataclasses import dataclass
import uuid
@dataclass
class SessionState:
"""Conversation state tied to a session ID."""
session_id: str
selected_metrics: list[str] = None
time_range: tuple[int, int] = None
selected_tags: dict = None
created_at: float = None
last_accessed_at: float = None
# In-memory session store (use Redis for distributed deployments)
_sessions: dict[str, SessionState] = {}
_session_timeout = 3600 # 1 hour
def get_or_create_session(session_id: str = None) -> SessionState:
"""Get existing session or create new one."""
if session_id and session_id in _sessions:
session = _sessions[session_id]
session.last_accessed_at = time.time()
return session
# Create new session
new_id = session_id or str(uuid.uuid4())
session = SessionState(
session_id=new_id,
selected_metrics=[],
time_range=(None, None),
selected_tags={},
created_at=time.time(),
last_accessed_at=time.time()
)
_sessions[new_id] = session
return session
@server.call_tool()
async def call_tool_handler(name: str, arguments: dict) -> list:
"""Tool handler with session state management."""
session_id = arguments.pop("session_id", None)
session = get_or_create_session(session_id)
if name == "select_metrics":
# Update session state
session.selected_metrics = arguments.get("metrics", [])
return [{"session_id": session.session_id, "status": "ok"}]
elif name == "select_time_range":
session.time_range = (
arguments.get("start_unix_ts"),
arguments.get("end_unix_ts")
)
return [{"session_id": session.session_id, "status": "ok"}]
elif name == "query_metrics":
# Use session state if not overridden
metrics = arguments.get("metrics") or session.selected_metrics
start, end = arguments.get("time_range") or session.time_range
results = await query_backend(
metrics=metrics,
start_ts=start,
end_ts=end
)
return [{"session_id": session.session_id, "data": results}]
else:
raise ValueError(f"Unknown tool: {name}")
# Background task: clean up expired sessions
async def cleanup_expired_sessions():
"""Remove sessions older than timeout."""
now = time.time()
expired = [
sid for sid, session in _sessions.items()
if now - session.last_accessed_at > _session_timeout
]
for sid in expired:
del _sessions[sid]
```
**Trade-offs:**
- ✅ Reduced latency: state is pre-computed and cached
- ✅ Cleaner UX: users don't repeat context
- ❌ State divergence: if client has stale session ID
- ❌ Memory overhead: session store grows with number of active conversations
---
### Pattern 4: Proxy Aggregator
**Purpose:** Fan out to multiple backend servers in parallel, merge results, return unified response.
**Real-world example:** Smartsheet + Airtable query aggregator — search for records in both systems, return combined results sorted by relevance.
**When to use:**
- You need to combine data from multiple backends
- Backends can be queried in parallel (no ordering dependency)
- Result merging is deterministic
**Implementation:**
```python
import asyncio
from typing import Any
@server.call_tool()
async def call_tool_handler(name: str, arguments: dict) -> list[dict]:
"""Fan-out to multiple backends, merge results."""
if name == "search_records":
query = arguments.get("query")
# Parallel queries to multiple backends
tasks = [
query_smartsheet(query),
query_airtable(query),
query_notion(query)
]
results = await asyncio.gather(*tasks, return_exceptions=True)
# Handle failures gracefully (partial results OK)
merged = []
for result in results:
if isinstance(result, Exception):
logger.error(f"Backend query failed: {result}")
continue # Skip failed backend
merged.extend(result)
# Sort merged results by relevance score
merged.sort(key=lambda x: x.get("relevance_score", 0), reverse=True)
return merged[:100] # Limit to top 100
else:
raise ValueError(f"Unknown tool: {name}")
async def query_smartsheet(query: str) -> list[dict]:
"""Query Smartsheet with timeout."""
try:
return await asyncio.wait_for(
smartsheet_client.search(query),
timeout=2.0
)
except asyncio.TimeoutError:
logger.warning("Smartsheet query timed out")
return []
async def query_airtable(query: str) -> list[dict]:
"""Query Airtable with timeout."""
try:
return await asyncio.wait_for(
airtable_client.search(query),
timeout=2.0
)
except asyncio.TimeoutError:
logger.warning("Airtable query timed out")
return []
async def query_notion(query: str) -> list[dict]:
"""Query Notion with timeout."""
try:
return await asyncio.wait_for(
notion_client.search(query),
timeout=2.0
)
except asyncio.TimeoutError:
logger.warning("Notion query timed out")
return []
```
**Trade-offs:**
- ✅ Unified query interface across multiple backends
- ✅ Parallel execution: total latency = max(backend latencies), not sum
- ❌ Result merging can be complex (ranking, deduplication)
- ❌ One slow backend affects overall latency
**Optimization:** Set per-backend timeout and include partial results from faster backends while slower ones still run.
---
### Pattern 5: Domain-Specific Adapter
**Purpose:** Implement specialized business logic on top of MCP. Example: Chronograph real-time time-series adapter.
**Real-world example:** Chronograph — expose time-series metrics as resources with automatic aggregation, downsampling, and real-time subscription support.
**When to use:**
- You need custom business logic beyond simple CRUD
- The adapter serves a specific use case (time-series, geospatial, graph queries)
- You want rich domain semantics in resource URIs
**Implementation:**
```python
from mcp.types import Resource, ReadResourceRequest
@server.list_resources()
async def list_resources_handler(request: ListResourcesRequest) -> list[Resource]:
"""Chronograph: Time-series metrics as resources."""
# Resource URI format: metric://namespace/metric_name?resolution=1m&aggregate=sum
resources = []
# Fetch available metrics from backend
for metric in await chronograph_backend.list_metrics():
# High-resolution resource (raw data)
resources.append(Resource(
uri=f"metric://{metric.namespace}/{metric.name}?resolution=raw",
name=f"{metric.name} (raw)",
description=f"Raw time-series data for {metric.name}",
mimeType="application/json"
))
# Downsampled resources (pre-aggregated for performance)
View on GitHub