| name | error-recovery-strategy |
| description | Classify error into 4 buckets (transient / deterministic / stale / unknown) and pick one of 5 actions (retry / switch / fallback / refresh-then-retry / ask-user / skip).
USE WHEN: tool returns non-success, sub-agent `status: closed-failed`, exception escapes, timeout fires, weird partial-success result, ECONNREFUSED / 5xx / 429 / timeout / permission denied / "command not found" / "fail" / "error" / "出错了" / "挂" / "失败".
TRIGGER PHRASES: "出错了", "failed", "挂", "error", "失败", "fail", "permission denied", "command not found", "ECONNREFUSED", "timeout", "挂了", "再试一次", "retry", "这不行", "没用", "fallback", "退路", "不行", "跑不通", "broken".
SKIP WHEN: operation succeeded, error is in user input (clarification case), error is part of expected flow (grep 0 matches).
|
| license | Apache-2.0 |
| compatibility | Requires MiniMax Code with Agent Plugins 1.0 support. |
| metadata | {"author":"antianqi","version":"0.1.2","inspired-by":"https://github.com/openai/codex/blob/main/codex-rs/code-mode/src/grpc_session/reconnect.rs and core/src/session/multi_agents.rs","changes-from-v0.1.1":"Audit gap: the sub-agent invocation example in '## Example' (line 115) used the Codex-style `task(subagent=...)` shape; switched to the canonical mcode 0.2.4 `task(agent_name=..., prompt=...)` shape (mcode accepts `subagent_type=` as a runtime alias but `agent_name=` is the canonical form per `cli.js:B6c`). The round-1 72952c9 amend touched 4 Skills (fork-context-decision / delegate-with-context / parallel-fanout / model-router) and missed this 5th; the v1.0.4 round-2 close-out also missed it. Caught by the v1.0.4 audit sweep across all 23 Skills' code blocks. The rest of the Skill body is unchanged."} |
Error Recovery Strategy
When something fails, the default human reaction is "retry." That is often the wrong
default. Retrying a permission-denied file write burns the same error three times in a row.
Retrying a network timeout that won't resolve in 30 seconds burns three minutes.
This Skill codifies the decision: categorize the error first, then pick one of five
recovery actions, then commit to it explicitly.
When to use
Activate when any of these is true:
- A tool call returns a non-success result (non-zero exit, HTTP 4xx/5xx, exception,
error message).
- A sub-agent reports
status: closed-failed in the family file.
- An exception escapes from any of your own code or a library you called.
- A timeout fires on a long-running operation.
- A "weird" result comes back that might be a partial success (e.g. command exited 0
but produced no output where you expected output).
When NOT to use
- The operation succeeded. Do not second-guess success.
- The error is in user input (bad prompt, missing file the user should provide). That is
not a recovery case; it is a clarification case.
- The error is part of expected flow (e.g. a
grep returning 0 matches is an exit-1, but
it is not a failure for the search use case).
Process
-
Stop. Do not retry yet. Even if the obvious answer is "retry," run this Skill.
-
Categorize the error into one of four buckets:
| Bucket | Signals | Examples |
|---|
| transient | Will probably succeed if tried again soon | Network timeout, HTTP 429/503, "ECONNRESET", "temporarily unavailable", rate limit |
| deterministic | Will fail every time the same way | Permission denied, file not found, "command not found", invalid syntax, wrong credentials |
| stale | Was true once, not anymore | "Index out of date", "stale lock", "file modified since read" |
| unknown | Cannot classify from the message alone | Generic exception, empty error, "something went wrong" |
-
Pick one of five recovery actions based on the bucket:
| Bucket | Default action | When to deviate |
|---|
| transient | retry-with-backoff (see next skill) | Never skip the backoff. If retries exceed the budget, escalate to ask-user. |
| deterministic | switch (try an alternative tool) or fallback (degrade gracefully) | If no alternative exists, ask-user. |
| stale | refresh-then-retry (re-read state, then retry once) | If still failing after one refresh, escalate to ask-user. |
| unknown | ask-user (or skip-with-warning if user signal allowed it) | Never silently retry an unknown error. |
-
Execute the chosen action explicitly — do not bundle it with other tool calls.
The user must be able to see "I retried" / "I switched to X" / "I asked you" as a discrete
step in the response.
-
Record the decision in the response so the user can audit:
Recovery decision: <action>
Reason: <one-line reason, citing the bucket>
-
After the action, re-evaluate. If the action also failed, re-categorize (the bucket
may have changed — e.g. a "transient" error that persists for 5 retries is now a
"deterministic" error in practice) and pick a new action.
Output contract
The user sees, in this order:
- One-line error summary (what went wrong, with the tool / file / op named).
- Bucket classification:
transient / deterministic / stale / unknown.
- Chosen action:
retry / switch / fallback / refresh-then-retry / ask-user / skip.
- One-line reason citing the bucket.
- (If retry) the backoff plan: "N attempts, M seconds between, exponential with jitter."
- (If ask) the user-facing question with concrete options.
- (If fallback / skip) what degraded behaviour the user is now getting.
Example
> bash(command="cat /etc/hosts")
Error: cat: /etc/hosts: Permission denied
Recovery decision: switch
Reason: bucket=deterministic (permission denied won't change on retry);
switching to `read` with explicit `path=/etc/hosts` (read may have different
permission semantics than shell cat) before asking the user.
> bash(command="curl https://internal-api.example.com/health")
Error: connect ECONNREFUSED 10.0.0.5:443
Recovery decision: retry
Reason: bucket=transient (network refused); will retry 3 times with 2s/4s/8s backoff
and 500ms jitter; if all fail, escalate to ask-user.
> task(agent_name="explore", prompt="...")
Status: closed-failed. Sub-agent error: "context window exceeded".
Recovery decision: ask-user
Reason: bucket=unknown (sub-agent did not return a clear error class); do not silently
retry with a smaller brief; surface to user with options:
(a) reduce the sub-task scope
(b) switch to a model with larger context
(c) skip this sub-task
Common pitfalls
- Do not default to retry. Retry is only correct for
transient (and a few stale).
For deterministic and unknown, retry is the most expensive wrong answer.
- Do not bundle the recovery with other tool calls. A retry hidden inside a larger
batch of work is invisible. Always surface the recovery as a discrete step.
- Do not re-categorize silently. If you categorize as
transient, retry 3 times,
and it still fails, the bucket is now deterministic or unknown — say so out loud.
- Do not ask the user a vague question. "What should I do?" is not an option. Give
the user 2-4 concrete options based on the bucket.
- Do not skip-with-warning without permission. The user did not pre-authorize
silent skips. If the work is optional, the user should have said so at the start.
- Do not blame the tool. The tool did what it was told. Categorize the error
honestly, not defensively.
- Do not loop on retry forever. Always have a max-attempt budget; on exhaustion,
escalate to
ask-user.
Verification checklist