Skip to main content

error-recovery-strategy

Classify error into 4 buckets (transient / deterministic / stale / unknown) and pick one of 5 actions (retry / switch / fallback / refresh-then-retry / ask-user / skip). USE WHEN: tool returns non-success, sub-agent `status: closed-failed`, exception escapes, timeout fires, weird partial-success result, ECONNREFUSED / 5xx / 429 / timeout / permission denied / "command not found" / "fail" / "error" / "出错了" / "挂" / "失败". TRIGGER PHRASES: "出错了", "failed", "挂", "error", "失败", "fail", "permission denied", "command not found", "ECONNREFUSED", "timeout", "挂了", "再试一次", "retry", "这不行", "没用", "fallback", "退路", "不行", "跑不通", "broken". SKIP WHEN: operation succeeded, error is in user input (clarification case), error is part of expected flow (grep 0 matches).

설치로 이동

소스 정보

저장소
MiniMax-AI/MiniMax-Code-Plugins
최근 소스 활동
2026년 9월 10일 01:39
감지된 SKILL.md 언어
영어
스타
11
포크
10

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.

SKILL.md 표시 중

SKILL.md
소스 지침 · 읽기 전용 미리보기
name
error-recovery-strategy
description
Classify error into 4 buckets (transient / deterministic / stale / unknown) and pick one of 5 actions (retry / switch / fallback / refresh-then-retry / ask-user / skip). USE WHEN: tool returns non-success, sub-agent `status: closed-failed`, exception escapes, timeout fires, weird partial-success result, ECONNREFUSED / 5xx / 429 / timeout / permission denied / "command not found" / "fail" / "error" / "出错了" / "挂" / "失败". TRIGGER PHRASES: "出错了", "failed", "挂", "error", "失败", "fail", "permission denied", "command not found", "ECONNREFUSED", "timeout", "挂了", "再试一次", "retry", "这不行", "没用", "fallback", "退路", "不行", "跑不通", "broken". SKIP WHEN: operation succeeded, error is in user input (clarification case), error is part of expected flow (grep 0 matches).
license
Apache-2.0
compatibility
Requires MiniMax Code with Agent Plugins 1.0 support.
metadata
{"author":"antianqi","version":"0.1.2","inspired-by":"https://github.com/openai/codex/blob/main/codex-rs/code-mode/src/grpc_session/reconnect.rs and core/src/session/multi_agents.rs","changes-from-v0.1.1":"Audit gap: the sub-agent invocation example in '## Example' (line 115) used the Codex-style `task(subagent=...)` shape; switched to the canonical mcode 0.2.4 `task(agent_name=..., prompt=...)` shape (mcode accepts `subagent_type=` as a runtime alias but `agent_name=` is the canonical form per `cli.js:B6c`). The round-1 72952c9 amend touched 4 Skills (fork-context-decision / delegate-with-context / parallel-fanout / model-router) and missed this 5th; the v1.0.4 round-2 close-out also missed it. Caught by the v1.0.4 audit sweep across all 23 Skills' code blocks. The rest of the Skill body is unchanged."}
# Error Recovery Strategy When something fails, the default human reaction is "retry." That is often the **wrong** default. Retrying a permission-denied file write burns the same error three times in a row. Retrying a network timeout that won't resolve in 30 seconds burns three minutes. This Skill codifies the decision: **categorize the error first, then pick one of five recovery actions, then commit to it explicitly.** ## When to use Activate when **any** of these is true: - A tool call returns a non-success result (non-zero exit, HTTP 4xx/5xx, exception, error message). - A sub-agent reports `status: closed-failed` in the family file. - An exception escapes from any of your own code or a library you called. - A timeout fires on a long-running operation. - A "weird" result comes back that might be a partial success (e.g. command exited 0 but produced no output where you expected output). ## When NOT to use - The operation succeeded. Do not second-guess success. - The error is in user input (bad prompt, missing file the user should provide). That is not a recovery case; it is a clarification case. - The error is part of expected flow (e.g. a `grep` returning 0 matches is an exit-1, but it is not a failure for the search use case). ## Process 1. **Stop. Do not retry yet.** Even if the obvious answer is "retry," run this Skill. 2. **Categorize the error** into one of four buckets: | Bucket | Signals | Examples | |---|---|---| | **transient** | Will probably succeed if tried again soon | Network timeout, HTTP 429/503, "ECONNRESET", "temporarily unavailable", rate limit | | **deterministic** | Will fail every time the same way | Permission denied, file not found, "command not found", invalid syntax, wrong credentials | | **stale** | Was true once, not anymore | "Index out of date", "stale lock", "file modified since read" | | **unknown** | Cannot classify from the message alone | Generic exception, empty error, "something went wrong" | 3. **Pick one of five recovery actions** based on the bucket: | Bucket | Default action | When to deviate | |---|---|---| | **transient** | `retry-with-backoff` (see next skill) | Never skip the backoff. If retries exceed the budget, escalate to `ask-user`. | | **deterministic** | `switch` (try an alternative tool) or `fallback` (degrade gracefully) | If no alternative exists, `ask-user`. | | **stale** | `refresh-then-retry` (re-read state, then retry once) | If still failing after one refresh, escalate to `ask-user`. | | **unknown** | `ask-user` (or `skip-with-warning` if user signal allowed it) | Never silently retry an unknown error. | 4. **Execute the chosen action explicitly** — do not bundle it with other tool calls. The user must be able to see "I retried" / "I switched to X" / "I asked you" as a discrete step in the response. 5. **Record the decision in the response** so the user can audit: ```text Recovery decision: <action> Reason: <one-line reason, citing the bucket> ``` 6. **After the action**, re-evaluate. If the action also failed, re-categorize (the bucket may have changed — e.g. a "transient" error that persists for 5 retries is now a "deterministic" error in practice) and pick a new action. ## Output contract The user sees, in this order: - One-line error summary (what went wrong, with the tool / file / op named). - Bucket classification: `transient` / `deterministic` / `stale` / `unknown`. - Chosen action: `retry` / `switch` / `fallback` / `refresh-then-retry` / `ask-user` / `skip`. - One-line reason citing the bucket. - (If retry) the backoff plan: "N attempts, M seconds between, exponential with jitter." - (If ask) the user-facing question with concrete options. - (If fallback / skip) what degraded behaviour the user is now getting. ## Example ```text > bash(command="cat /etc/hosts") Error: cat: /etc/hosts: Permission denied Recovery decision: switch Reason: bucket=deterministic (permission denied won't change on retry); switching to `read` with explicit `path=/etc/hosts` (read may have different permission semantics than shell cat) before asking the user. ``` ```text > bash(command="curl https://internal-api.example.com/health") Error: connect ECONNREFUSED 10.0.0.5:443 Recovery decision: retry Reason: bucket=transient (network refused); will retry 3 times with 2s/4s/8s backoff and 500ms jitter; if all fail, escalate to ask-user. ``` ```text > task(agent_name="explore", prompt="...") Status: closed-failed. Sub-agent error: "context window exceeded". Recovery decision: ask-user Reason: bucket=unknown (sub-agent did not return a clear error class); do not silently retry with a smaller brief; surface to user with options: (a) reduce the sub-task scope (b) switch to a model with larger context (c) skip this sub-task ``` ## Common pitfalls - **Do not default to retry.** Retry is only correct for `transient` (and a few `stale`). For `deterministic` and `unknown`, retry is the most expensive wrong answer. - **Do not bundle the recovery with other tool calls.** A retry hidden inside a larger batch of work is invisible. Always surface the recovery as a discrete step. - **Do not re-categorize silently.** If you categorize as `transient`, retry 3 times, and it still fails, the bucket is now `deterministic` or `unknown` — say so out loud. - **Do not ask the user a vague question.** "What should I do?" is not an option. Give the user 2-4 concrete options based on the bucket. - **Do not skip-with-warning without permission.** The user did not pre-authorize silent skips. If the work is optional, the user should have said so at the start. - **Do not blame the tool.** The tool did what it was told. Categorize the error honestly, not defensively. - **Do not loop on retry forever.** Always have a max-attempt budget; on exhaustion, escalate to `ask-user`. ## Verification checklist - [ ] Did you categorize the error into one of four buckets before picking an action? - [ ] Did you pick one of five actions based on the bucket (not the default)? - [ ] Did you state the recovery decision in the response, with the bucket and reason? - [ ] (Retry) Did you specify the backoff plan (attempts, intervals, jitter)? - [ ] (Ask) Did you give 2-4 concrete options, not "what should I do?" - [ ] (Switch / Fallback) Did you name the alternative tool / the degraded behaviour? - [ ] (Skip) Did you confirm the user pre-authorized this work as optional? - [ ] Did you re-evaluate after the action and re-categorize if it failed? - [ ] Is the recovery step a discrete line in the response (not bundled)?
GitHub에서 보기