| name | u0149-tooling-disaster-recovery-orchestrator |
| description | Build and operate the "Tooling Disaster Recovery Orchestrator" capability for Tool Reliability and Execution Quality. Trigger when this exact capability is needed in mission execution. |
Tooling Disaster Recovery Orchestrator
Why This Skill Exists
We need this skill because automation collapses when tools are flaky and failure modes are opaque. This specific skill improves recovery speed after severe outages.
When To Use
Use this skill when the request explicitly needs "Tooling Disaster Recovery Orchestrator" outcomes in the Tool Reliability and Execution Quality domain.
Step-by-Step Implementation Guide
- Define the scope and success metrics for
Tooling Disaster Recovery Orchestrator, including at least three measurable KPIs tied to silent failures and cascading retries.
- Design and version the input/output contract for tool runs, error signatures, and retry outcomes, then add schema validation and failure-mode handling.
- Implement the core capability using failover and restoration sequencing, and produce recovery mission plans with deterministic scoring.
- Integrate the skill into swarm orchestration: task routing, approval gates, retry strategy, and rollback controls.
- Add unit, integration, and simulation tests that explicitly cover silent failures and cascading retries, then run regression baselines.
- Deploy behind a feature flag, monitor telemetry/alerts for two release cycles, and iterate thresholds based on observed outcomes.
Required Deliverables
- Capability contract: input schema, deterministic scoring, output schema, and failure modes.
- Runtime profile: general-capability using failover and restoration sequencing to produce recovery mission plans.
- Orchestration integration: tool-reliability-and-execution-quality:general-capability routing, approval gates, retries, and rollback controls.
- Validation evidence: unit, integration, simulation, regression-baseline suites and rollout telemetry.
Operational Runbook
Preflight
- Confirm the Tooling Disaster Recovery Orchestrator request scope, source evidence, and measurable success criteria before execution.
- Verify feature flag skill_0149_tooling-disaster-recovery-orches, approval gates, and rollback owner before autonomous use.
Execution
- Execute failover and restoration sequencing with deterministic scoring and reproducible trace capture.
- Produce recovery mission plans plus scorecard, assumptions, and unresolved-risk notes.
Recovery
- Fail closed when required signals, evidence, or approval gates are missing.
- Rollback to the last stable baseline when posture is critical or validation fails.
Handoff
- Publish recovery mission plans, validation evidence, and telemetry links to downstream owners.
- Queue follow-up tasks for unresolved risks, threshold tuning, or approval review.
Guardrails
- [quality] Require deterministic scoring and validation evidence before promotion.
- [reliability] Preserve retries, rollback controls, and failure-mode evidence for every run.
- [safety] Route critical posture or missing approval gates to human review before autonomous action.