- name
- error-recovery
- description
- Handle errors, timeouts, and failures in multi-agent workflows. Use when dealing with external model timeouts, API failures, partial success, user cancellation, or graceful degradation.
- user-invocable
- false
# Error Recovery
**Version:** 1.2.0
**Purpose:** Patterns for handling failures in multi-agent workflows
**Status:** Production Ready
## Overview
Error recovery is the practice of handling failures gracefully in multi-agent workflows, ensuring that temporary errors, timeouts, or partial failures don't derail entire workflows. In production systems with external dependencies (AI models, APIs, network calls), failures are inevitable. The question is not "will it fail?" but "how will we handle it when it does?"
This skill provides battle-tested patterns for:
- **User escalation** (STOP and report before fallback — DEFAULT)
- **Timeout handling** (external models taking >30s)
- **API failure recovery** (401, 500, network errors)
- **Partial success strategies** (some agents succeed, others fail)
- **User cancellation** (graceful Ctrl+C handling)
- **Missing tools** (claudish not installed)
- **Out of credits** (payment/quota errors)
- **Retry strategies** (exponential backoff, max retries)
With proper error recovery, workflows become **resilient** and **production-ready**.
## Core Patterns
### Pattern 0: User Escalation (DEFAULT — Read First)
**This is the MOST IMPORTANT pattern. It overrides all other patterns when the user has requested a specific model.**
**Rule: NEVER silently substitute, retry with a different model, or fall back to embedded Claude when the user requested a specific model. STOP and REPORT the failure first.**
**When this applies:**
- User explicitly requested a model (e.g., "use Gemini to redesign X")
- User selected specific models for a task
- Any workflow where the model choice was intentional
**When graceful degradation (Patterns 1-7) applies instead:**
- Automated pipelines where *any result* > *no result*
- User explicitly said "use whatever works"
- `/team` workflows where the command already handles failure reporting
**The Protocol:**
```
Step 1: Model fails (non-zero exit, empty output, API error, rate limit, binary crash)
Step 2: STOP immediately. Do NOT:
❌ Silently launch a different model
❌ Retry with a different provider prefix
❌ Fall back to embedded Claude without asking
❌ Run `claudish --top-models` or invent model IDs — use the live catalog (`list_models`) instead
❌ Substitute a "similar" model
Step 3: REPORT to the user with:
- What model was requested
- What happened (exact error message)
- How many attempts were made and what was tried
- Actionable options for the user to choose from
Step 4: WAIT for user's decision before proceeding
```
**Report Template:**
```
"{Model Name} failed — {error category}.
What happened:
1. Attempt 1: {what was tried} — {exact error}
2. Attempt 2: {what was tried} — {exact error}
Options:
(1) {Fix and retry} — {specific fix description}
(2) Use a different model — {suggest alternatives if known}
(3) Skip this model and continue without it
(4) Cancel the workflow
(5) Report this error to claudish developers
Which do you prefer?"
```
**Why this matters:** When a user says "use Gemini", they've made a deliberate choice — for its 1M context, its reasoning style, or for model diversity. Silently substituting GPT-5 defeats the purpose. The user should always be in control of model selection decisions.
---
### Pattern 1: Timeout Handling
**Scenario: External Model Takes >30s**
External AI models via Claudish may take >30s due to:
- Model service overloaded (high demand)
- Network latency (slow connection)
- Complex task (large input, detailed analysis)
- Model thinking time (GPT-5, Grok reasoning models)
**Detection:**
```
Monitor execution time and set timeout limits:
const TIMEOUT_THRESHOLD = 30000; // 30 seconds
startTime = Date.now();
executeClaudish(model, prompt);
setInterval(() => {
elapsedTime = Date.now() - startTime;
if (elapsedTime > TIMEOUT_THRESHOLD && !modelResponded) {
handleTimeout();
}
}, 1000);
```
**Recovery Strategy:**
```
Step 1: Detect Timeout
Log: "Timeout: grok after 30s with no response"
Step 2: Notify User
Present options:
"Model 'Grok' timed out after 30 seconds.
Options:
1. Retry with 60s timeout
2. Skip this model and continue with others
3. Cancel entire workflow
What would you like to do? (1/2/3)"
Step 3a: User selects RETRY
Increase timeout to 60s
Re-execute claudish with longer timeout
If still times out: Offer skip or cancel
Step 3b: User selects SKIP
Log: "Skipping Grok review due to timeout"
Mark this model as failed
Continue with remaining models
(Graceful degradation pattern)
Step 3c: User selects CANCEL
Exit workflow gracefully
Save partial results (if any)
Log cancellation reason
```
**Graceful Degradation:**
```
Multi-Model Review Example:
Requested: 5 models (Claude, Grok, Gemini, GPT-5, DeepSeek)
Timeout: Grok after 30s
Result:
- Claude: Success ✓
- Grok: Timeout ✗ (skipped)
- Gemini: Success ✓
- GPT-5: Success ✓
- DeepSeek: Success ✓
Successful: 4/5 models (80%)
Threshold: N ≥ 2 for consolidation ✓
Action:
Proceed with consolidation using 4 reviews
Notify user: "4/5 models completed (Grok timeout). Proceeding with 4-model consensus."
Benefits:
- Workflow completes despite failure
- User gets results (4 models better than 1)
- Timeout doesn't derail entire workflow
```
**Example Implementation:**
```
# Via create_session MCP tool (timeout handled by the tool)
create_session(model="grok", prompt=PROMPT, timeout_seconds=30)
# React to channel events:
# - completed → get_output(session_id) → process result
# - failed → inspect error content:
# - Content contains "timeout" → timeout occurred
# - Content contains "401" → API key issue
# - Other → general failure
# Via team MCP tool. There is NO per-model timeout any more — the parameter was
# removed, and passing one is silently ignored. You bound the wait yourself by
# bounding the poll loop.
team(mode="run", path=SESSION_DIR, models=["grok"], input_file=INPUT_MD,
require_pattern=<regex for the shape PROMPT mandates>)
team(mode="status", path=SESSION_DIR) # poll until no slot has state === "RUNNING"
# Check per-model status in the SETTLED status response, not the run response — run
# returns before any model has answered. A slot reported EMPTY with reason
# shape_mismatch is a FAILURE to recover from, not a short answer to accept — the
# model finished without producing the required shape.
#
# Before treating a quiet slot as hung, read idle_seconds_by_slot together with
# activity_by_slot: 90s idle in "tool_executing" is a build or test suite and is
# normal; 90s idle in "running" is a model that stopped mid-answer.
```
---
### Pattern 2: API Failure Recovery
**Common API Failure Scenarios:**
```
401 Unauthorized:
- Invalid API key (OPENROUTER_API_KEY incorrect)
- Expired API key
- API key not set in environment
500 Internal Server Error:
- Model service temporarily down
- Server overload
- Model deployment issue
Network Errors:
- Connection timeout (network slow/unstable)
- DNS resolution failure
- Firewall blocking request
429 Too Many Requests:
- Rate limit exceeded
- Too many concurrent requests
- Quota exhausted for time window
```
**Recovery Strategies by Error Type:**
**401 Unauthorized:**
> **Pattern 0 guard:** If the user requested a specific model, apply Pattern 0 (stop and report with options) instead of auto-fallback. The fallback below applies only to automated pipelines or when the user pre-authorized graceful degradation.
```
Detection:
API returns 401 status code
Recovery:
1. Log: "API authentication failed (401)"
2. Check if OPENROUTER_API_KEY is set:
if [ -z "$OPENROUTER_API_KEY" ]; then
notifyUser("OpenRouter API key not found. Set OPENROUTER_API_KEY in .env")
else
notifyUser("Invalid OpenRouter API key. Check .env file")
fi
3. Skip all external models
4. Fallback to embedded Claude only
5. Notify user:
"⚠️ API authentication failed. Falling back to embedded Claude.
To fix: Add valid OPENROUTER_API_KEY to .env file."
No retry (authentication won't fix itself)
```
**500 Internal Server Error:**
```
Detection:
API returns 500 status code
Recovery:
1. Log: "Model service error (500): grok"
2. Wait 5 seconds (give service time to recover)
3. Retry ONCE
4. If retry succeeds: Continue normally
5. If retry fails: Skip this model, continue with others
Example:
try {
result = await claudish(model, prompt);
} catch (error) {
if (error.status === 500) {
log("500 error, waiting 5s before retry...");
await sleep(5000);
try {
result = await claudish(model, prompt); // Retry
log("Retry succeeded");
} catch (retryError) {
log("Retry failed, skipping model");
skipModel(model);
continueWithRemaining();
}
}
}
Max retries: 1 (avoid long delays)
```
**Network Errors:**
```
Detection:
- Connection timeout
- ECONNREFUSED
- ETIMEDOUT
- DNS resolution failure
Recovery:
Retry up to 3 times with exponential backoff:
async function retryWithBackoff(fn, maxRetries = 3) {
for (let i = 0; i < maxRetries; i++) {
try {
return await fn();
} catch (error) {
if (!isNetworkError(error)) throw error; // Not retriable
if (i === maxRetries - 1) throw error; // Max retries reached
const delay = Math.pow(2, i) * 1000; // 1s, 2s, 4s
log(`Network error, retrying in ${delay}ms (attempt ${i+1}/${maxRetries})`);
await sleep(delay);
}
}
}
result = await retryWithBackoff(() => claudish(model, prompt));
Rationale: Network errors are often transient (temporary)
```
**429 Rate Limiting:**
```
Detection:
API returns 429 status code
Response may include Retry-After header
Recovery:
1. Check Retry-After header (seconds to wait)
2. If present: Wait for specified time
3. If not present: Wait 60s (default)
4. Retry ONCE after waiting
5. If still rate limited: Skip model
Example:
if (error.status === 429) {
const retryAfter = error.headers['retry-after'] || 60;
log(`Rate limited. Waiting ${retryAfter}s before retry...`);
await sleep(retryAfter * 1000);
try {
result = await claudish(model, prompt);
} catch (retryError) {
log("Still rate limited after retry. Skipping model.");
skipModel(model);
}
}
Note: Respect Retry-After header (avoid hammering API)
```
**Graceful Degradation for All API Failures:**
> **Pattern 0 guard:** If the user requested specific models, apply Pattern 0 (stop and report with options) instead of auto-fallback. The fallback below applies only to automated pipelines or when the user pre-authorized graceful degradation.
```
Fallback Strategy:
If ALL external models fail (401, 500, network, etc.):
1. Log all failures
2. Notify user:
"⚠️ All external models failed. Falling back to embedded Claude.
Errors:
- Grok: Network timeout
- Gemini: 500 Internal Server Error
- GPT-5: Rate limited (429)
- DeepSeek: Authentication failed (401)
Proceeding with Claude Sonnet (embedded) only."
3. Run embedded Claude review
4. Present results with disclaimer:
"Review completed using Claude only (external models unavailable).
For multi-model consensus, try again later."
Benefits:
- User still gets results (better than nothing)
- Workflow completes (not aborted)
- Clear error communication (user knows what happened)
```
---
### Pattern 3: Partial Success Strategies
**Scenario: 2 of 4 Models Complete Successfully**
In multi-model workflows, it's common for some models to succeed while others fail.
**Tracking Success/Failure:**
```
const results = await Promise.allSettled([
Agent({ subagent: "reviewer", model: "claude" }),
Agent({ subagent: "reviewer", model: "grok" }),
Agent({ subagent: "reviewer", model: "gemini" }),
Agent({ subagent: "reviewer", model: "gpt-5" })
]);
const successful = results.filter(r => r.status === 'fulfilled');
const failed = results.filter(r => r.status === 'rejected');
log(`Success: ${successful.length}/4`);
log(`Failed: ${failed.length}/4`);
```
**Decision Logic:**
```
If N ≥ 2 successful:
→ Proceed with consolidation
→ Use N reviews (not all 4)
→ Notify user about failures
If N < 2 successful:
→ Insufficient data for consensus
→ Offer user choice:
1. Retry failures
2. Abort workflow
3. Proceed with embedded Claude only
Example:
successful.length = 2 (Claude, Gemini)
failed.length = 2 (Grok timeout, GPT-5 500 error)
Action:
notifyUser("2/4 models completed successfully. Proceeding with consolidation using 2 reviews.");
consolidateReviews([
"ai-docs/claude-review.md",
"ai-docs/gemini-review.md"
]);
presentResults({
totalModels: 4,
successful: 2,
failureReasons: {
grok: "Timeout after 30s",
gpt5: "500 Internal Server Error"
}
});
```
**Communication Strategy:**
```
Be transparent with user about partial success:
❌ WRONG:
"Multi-model review complete!"
(User assumes all 4 models ran)
✅ CORRECT:
"Multi-model review complete (2/4 models succeeded).
Successful:
- Claude Sonnet ✓
- Gemini 2.5 Flash ✓
Failed:
- Grok: Timeout after 30s
- GPT-5 Codex: 500 Internal Server Error
Proceeding with 2-model consensus.
Top issues: [...]"
User knows:
- What succeeded (Claude, Gemini)
- What failed (Grok, GPT-5)
- Why they failed (timeout, 500 error)
- What action was taken (2-model consensus)
```
**Consolidation Adapts to N Models:**
```
Consolidation logic must handle variable N:
✅ CORRECT - Flexible N:
function consolidateReviews(reviewFiles) {
const N = reviewFiles.length;
log(`Consolidating ${N} reviews`);
// Consensus thresholds adapt to N
const unanimousThreshold = N; // All N agree
const strongThreshold = Math.ceil(N * 0.67); // 67%+ agree
const majorityThreshold = Math.ceil(N * 0.5); // 50%+ agree
// Apply consensus analysis with dynamic thresholds
...
}
❌ WRONG - Hardcoded N:
// Assumes always 4 models
const unanimousThreshold = 4; // Breaks if N = 2!
```
---
### Pattern 4: User Cancellation Handling (Ctrl+C)
**Scenario: User Presses Ctrl+C During Workflow**
Users may cancel long-running workflows for various reasons:
- Taking too long
- Realized they want different configuration
- Accidentally triggered workflow
- Need to prioritize other work
**Cleanup Strategy:**
```
process.on('SIGINT', async () => {
log("⚠️ User cancelled workflow (Ctrl+C)");
// Step 1: Stop all running processes gracefully
await stopAllAgents();
Ver en GitHub