| name | operational-resilience-tester |
| description | Review backup and restore, failover, disaster recovery, restart procedures, crisis exercises, scenario tests, and lessons learned. |
| version | 1.1.0 |
| since | 2026-06-17 |
| last_modified | 2026-06-17 |
| authors | ["platform-engineering"] |
| stability | stable |
| min_platform_version | {"codex":"unknown","amazon-q":"unknown","antigravity":"unknown","auggie":"unknown","bob":"unknown","claude-code":"unknown","cline":"unknown","codebuddy":"unknown","continue":"unknown","costrict":"unknown","crush":"unknown","github-copilot":"unknown","gitlab-duo":"unknown","factory":"unknown","forgecode":"unknown","opencode":"unknown","openhands":"unknown","cursor":"unknown","roo-code":"unknown","kiro":"unknown","junie":"unknown","gemini-cli":"unknown","iflow":"unknown","kilocode":"unknown","kimi":"unknown","lingma":"unknown","pi":"unknown","qoder":"unknown","qwen":"unknown","windsurf":"unknown","ollama":"unknown"} |
| deprecated_since | null |
| replaces | null |
| supersedes | [] |
| changelog | [{"version":"1.1.0","date":"2026-06-17","change":"Initial generated production-ready SDLC / DevSecOps skill"}] |
Operational Resilience Tester
Purpose
Review backup and restore, failover, disaster recovery, restart procedures, crisis exercises, scenario tests, and lessons learned. Treat regulatory, security, and operational references as review and evidence guidance, not legal advice.
When to use
- operational resilience testing decisions, controls, or operating practices need independent review.
- A change affects operational resilience testing artifacts such as backup policy, restore test log, failover runbook, DR plan, crisis exercise report, lessons-learned register.
- The user needs evidence-oriented findings for risks such as unproven restore, manual failover bottleneck, RTO mismatch, unrehearsed crisis role, scenario gap, unclosed lesson.
- Audit, security, operations, or platform stakeholders need a concise readiness position.
- Existing documentation, tickets, tests, or logs must be turned into actionable remediation items.
Operating model
- Identify the relevant operational resilience testing artifacts, owners, systems, environments, and review boundary.
- Compare the available artifacts against expected signals such as RPO evidence, RTO measurement, failover transcript, exercise participant record, defect ticket, retest proof.
- Separate confirmed gaps from assumptions, missing evidence, and advisory improvement opportunities.
- Rate findings by operational, security, compliance, customer, and auditability impact.
- Recommend minimal remediation steps, validation evidence, owners, and review cadence.
Spec-Driven Change Context
- Treat repository specs, ADRs, runbooks, change proposals, design notes, and task files as durable context that outlives a chat session.
- For non-trivial changes, prefer a checked-in change artifact or equivalent proposal/design/tasks record before implementation begins.
- Capture requirement deltas explicitly: added, modified, removed, deprecated, or unchanged behavior.
- Keep implementation tasks traceable to acceptance criteria, affected specs, validation commands, and owners.
- During verification, compare the implementation against the proposal, design decisions, task checklist, and spec deltas.
- After completion, sync or archive completed change artifacts so the repository's source of truth reflects the final behavior.
- If the repository has no spec workflow yet, report the missing artifact and provide a minimal proposal/spec/tasks outline instead of relying on chat-only intent.
Skill-Specific Review Scope
- Primary artifacts: backup policy, restore test log, failover runbook, DR plan, crisis exercise report, lessons-learned register.
- Risk themes: unproven restore, manual failover bottleneck, RTO mismatch, unrehearsed crisis role, scenario gap, unclosed lesson.