| name | service-failure-recovery |
| description | Design service recovery for failures involving multiple teams/systems by defining detection, ownership, customer communication, compensation/remediation boundaries, escalation, and closure evidence. |
Service Failure Recovery
Use when this procedure is the primary professional method needed for the assignment.
Procedure
- Confirm the decision or outcome this work must support, its scope, owner, constraints, and definition of success.
- Establish the evidence baseline using incident/support data, policies, system capabilities, authority boundaries, and customer expectations. Do not fill material gaps with assumptions when they can change the result.
- Classify failure scenarios and customer impact, define first owner, preserve context across handoff, create recovery paths and exception authority, and close loop with learning.
- Exercise realistic edge, failure, transition, or exception cases that could invalidate the result; record unresolved uncertainty explicitly.
- Validate the output against the original outcome and any neighboring professional contracts so this skill does not silently absorb another specialist's authority.
- Record the resulting artifact, measurements, decisions, provenance, and handoff information needed for another owner to reproduce or continue the work.
Quality gate
A failed service has one understandable recovery path and customers are not bounced between owners without context.