Use when producing or updating AI-agent operations runbook for agent health, pauses, retries, replay, tool failures, containment, and operator authority. Use incident-response-runbook for the neighbouring concern; this skill owns the named document contract and its acceptance evidence.
Instalar com Codex ou Claude Copie este prompt, cole no Codex, Claude ou outro assistente e deixe que ele revise a página da skill e instale para você.
Um comando direto ignora o prompt de revisão. Verifique a origem antes de executá-lo.
Instruções da origem · Visualização somente leitura
name
14-ai-agent-runbook
description
Use when producing or updating AI-agent operations runbook for agent health, pauses, retries, replay, tool failures, containment, and operator authority. Use incident-response-runbook for the neighbouring concern; this skill owns the named document contract and its acceptance evidence.
Failures and unavailable checks cannot appear as passes.
Review record
Reviewer, date, disposition, open actions
The consumer can reproduce the acceptance decision.
Capability and Permission Boundaries
Minimum capabilities: read and search the authorised project sources. Execution is optional and limited to non-destructive validation.
Inspection is read-only by default. Create or edit the named project document only when explicitly authorised. Production mutation, publishing, destructive action, spending, external communication, or certification claims require separate explicit authority.
Treat secrets, tenant data, incident evidence, and financial records as least-privilege inputs; expose only the minimum evidence needed for review.
Degraded Mode
If files, execution, network, rendering, environment access, fonts, or current evidence are unavailable, return the narrowest useful draft plus a gap register. Label affected checks not assessed, retain the intended acceptance oracle, and state who must supply or verify the missing evidence. Never convert an unavailable check into a pass.
Decision Rules
Choice
Action
Failure or risk avoided
Evidence is complete and authority is explicit
Choose operation from run state, tool reversibility, and incident severity and produce the full artefact.
Duplicate, unauthorised, or unreplayable actions.
A required source or approval is missing
Stop the affected branch; record the gap, owner, and unblock condition.
Fabricated requirements or unauthorised action.
Evidence conflicts across sources
Preserve both claims, identify the controlling owner, and request a recorded decision.
Silent selection of a convenient but wrong source.
A check cannot run in the available environment
Keep its oracle and mark it not assessed; require later execution evidence.
False assurance from capability limits.
Workflow
Confirm the named deliverable, consumer, scope, environment, authority, and neighbouring-skill boundary.
Inventory required sources and validate provenance, freshness, internal consistency, and missing inputs. Stop the affected branch on a mandatory gap.
Extract traceable requirements, invariants, risks, and measurable acceptance criteria; record conflicts before choosing a design or procedure.
Apply the decision rules and the domain workflow below. For a failed branch, preserve evidence, choose the documented recovery path, or escalate to the named owner.
Draft the artefact, decision register, and evidence record together. Do not defer failure handling, rollback, security, tenancy, accessibility, or operational ownership.
Run available checks, review every result, repair failures, and hand off only when acceptance is observable. If recovery fails or authority is exceeded, stop and escalate without mutation.
Quality Standards
Ground every section in a named project source, decision, measured result, or accountable owner.
Give each requirement or procedure a deterministic oracle that another reviewer can reproduce.
Keep assumptions, exclusions, degraded checks, residual risks, and waivers visible at handoff.
Preserve the domain invariants and more specific controls in the existing workflow below; this contract does not replace them.
Run the repository anti-AI-slop gate: remove filler, verify named standards and dependencies, and retain purposeful domain detail.
Anti-Patterns
Copying a generic template without mapping it to project sources. Fix: attach each section to an approved requirement, configuration, risk, or owner.
Choosing a threshold because it is common practice. Fix: derive it from a requirement, measured baseline, risk decision, or current verified source.
Reporting an inaccessible or unexecuted check as passed. Fix: mark it not assessed, preserve the oracle, and name the verifier.
Mixing the neighbouring incident-response-runbook concern into this artefact without a boundary. Fix: cross-reference its output and keep ownership explicit.
Omitting failure, rollback, empty-state, security, tenancy, or escalation behaviour. Fix: specify the trigger, safe action, verification, and owner for each applicable case.
Mutating a repository, environment, tenant, ledger, or external system while drafting guidance. Fix: remain read-only until the exact mutation and authority are explicit.
Claiming compliance, certification, readiness, or release from prose alone. Fix: require source-attributed evidence and a named acceptance decision.
Worked Example
Given an approved project source and a conflicting implementation detail, record both with provenance, stop the affected branch, and obtain the accountable owner's decision. Then update the relevant contract, define a reproducible acceptance check, and retain its observed result. The artefact is accepted only when every operator action has an authority check, observable precondition, safe execution step, verification, and audit evidence.
References
logic.prompt - load only when its template, logic, or detail is needed.
README.md - load only when its template, logic, or detail is needed.
Core Instructions
Step 1: Kill-switch operations
Document three switches with operator-only access:
Global kill-switch — refuses every tool with kill_switch.global = refuse across every tenant. Two-person rule. Used for upstream provider compromise, mass irreversible-incident, or red-team-confirmed catastrophic vulnerability.
Per-tenant kill-switch — refuses tools for one tenant. Used for tenant-scoped incident, contractual obligation, or customer request.
Per-feature kill-switch — refuses tools for one agent feature globally. Used for feature-scoped regression or SEV1 from the SLO doc.
State the propagation SLA (default 5 s), the rehearsal cadence (monthly in staging), and the operator surface (ops console + API + on-call paging integration).
Step 2: Force-pause and force-resume
Force-pause an agent run — orchestrator marks the run intervened; in-flight tool calls complete or time-out; no further steps.
Force-pause all runs for a tenant — soft kill-switch; existing runs paused; new starts refused.
Force-resume — orchestrator resumes from last durable state with reasoning logged; available only after a SEV closure review.
Step 3: Replay a run
Operator can replay any historical run against the current planner + catalogue in the eval environment. Used for postmortems and for "would the fix have caught this?" verification.
Step 4: Agent-task quarantine
When a run is suspected harmful but not actioned, quarantine the run:
Mark the run quarantined.
All tool calls refused.
Run preserved with full state for forensic review.
Notify the tenant admin within 24 h.
Step 5: Audit-log review cadence
Daily: irreversible-action audit log reviewed by an on-call operator; any anomaly creates a ticket.
Weekly: per-feature audit summary reviewed by the AI lead.
Quarterly: tenant-facing audit export offered to Enterprise tenants.
Step 6: Agent-incident playbooks
Define playbooks for:
Mass irreversible-action incident.
Cross-tenant tool-routing attempt detected.
Indirect prompt injection succeeded in production.
Every drill produces evidence consumed by the SOC 2 / ISO / HIPAA control packs. Capture format and cadence are defined in 09-governance-compliance/25-ai-agent-evidence-pack-spec/references/ai-agent-evidence-frequency-table.md (rows 15, 16, 17):
Drill
Evidence
Cadence
Controls
Kill-switch (global / per-tenant / per-feature)
drill report; audit-log entry; propagation timing
quarterly staging; annual production
SOC2 CC7.4, A1.3; ISO A.5.30, A.8.2; HIPAA 164.308.a.7
Replay-a-run
drill report
quarterly
SOC2 A1.3; ISO A.5.30
Force-pause + force-resume
drill report
quarterly
SOC2 A1.3; ISO A.5.30
Agent-task quarantine
drill report; tenant-admin notification log
annual
SOC2 CC7.3; ISO A.5.25
The compliance runbook (06-deployment-operations/20-ai-agent-compliance-runbook) sets the calendar.