| name | resilience-audit |
| description | Failure-mode audit (FMEA for software) โ for each way the system can fail (network, storage, partial completion, crash, concurrency, bad input), check whether code DETECTS, HANDLES, RECOVERS, and COMMUNICATES it. Triggers on: "/resilience-audit", "resilience-audit", "FMEA audit". Use when touching network, storage, async, retry, or rollback paths. Flags data loss, silent-success-on-failure, missing rollback/retry/idempotency. Reports; does not fix unless asked. |
Resilience Audit
For every operation: "what happens when this FAILS?" Report; do NOT fix unless asked.
Failure categories
- External I/O โ network down/slow, API 4xx/5xx/timeout, rate-limit. Retry w/ backoff? Timeout set? Clear error vs hang?
- Storage โ disk full, permission denied, partial write. Atomic write (temp+rename)? Cleanup on failure? Existing good copy untouched?
- Partial completion โ half-done op (50/100 files). Reported as FAILURE, never success.
- Crash / OOM โ killed mid-op. Idempotent restart? No orphaned half-state?
- Concurrency โ two instances, race, deadlock. Locking / idempotency / safe re-entry?
- Input / data โ malformed, null, truncated, huge. Validate at boundary? Fail-fast?
- Dependency down โ fallback/cache/graceful degrade? Clear error vs silent hang?
- Resource exhaustion โ bounded? Backpressure? Cleanup on error path?
Per-stack timeout/atomicity/idempotency patterns to grep: read references/checks.md before scanning.
For each failure point, check 4 things
- Detected? code notices it (doesn't swallow)?
- Handled? retry/fallback/fail-clean โ not ignored, not silent-success?
- Recoverable? rollback/idempotent; no data loss or corruption?
- Communicated? clear error to user+log; not a hang, not a false "done"?
Discipline
- Trace actual failure path (cite file:line). Don't assume handling exists; prove it.
- "partial = failure" โ any path reporting success on partial completion = CRITICAL.
- "logged" โ "handled" โ swallowed+logged error that corrupts state or returns success = CRITICAL.
Fix mode (choice-gated)
After the report, present via ask_question:
- Fix safe ones โ add missing timeout, null/input validation, clear error+log on unhandled path. Each: checkpoint โ fix โ build+tests โ revert if newly red.
- Let me pick โ user-selected fixes only.
- Report only โ change nothing.
NEVER auto-fix: retry/rollback/recovery/atomicity logic (semantic changes can introduce new failure modes).
Output
| operation | failure mode | effect | handling (file:line) | severity | recommended guard |
Ordering/atomicity findings ยท Summary (counts + top fixes) ยท Not assessed
Severity: CRITICAL (data loss/corruption/silent-success) ยท HIGH (crash/hang/partial-no-recovery) ยท MEDIUM (poor degradation/missing retry) ยท LOW (cosmetic)