ワンクリックで
ops-fires
Production incidents dashboard. Reads ECS health, Sentry errors, CI failures. Offers to dispatch fix agents for active fires.
Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。
メニュー
Production incidents dashboard. Reads ECS health, Sentry errors, CI failures. Offers to dispatch fix agents for active fires.
Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。
Interactive OAuth init wizard for the multi-account Claude rotator. Walks through every account in the rotation config and, for any account missing a valid keychain token, delegates to the proven `rotate.mjs` magic-link flow (browser-driver cascade + Gmail polling), which writes the verified OAuth token to `Claude-Rotation-<key>` (key = account label or email, keychain account `$USER`). Re-runnable any time. Standalone alias of the same step inside `/ops:setup`.
Multi-account Claude Max rotator. Status, manual rotation, account list, add-account wizard, and CRS relay-pool auto-prioritization. Requires account_rotation_enabled=true in plugin settings.
Full inbox management across all channels — WhatsApp (whatsmeow bridge via mcp__whatsapp__*), iMessage (chat.db reader + AppleScript send via mcp__plugin_imessage_imessage__*), Email (Gmail MCP), Slack (MCP), Telegram (user-auth MCP), Discord (webhook + REST read), Notion (MCP — comments, mentions, assigned tasks). Scans FULL inbox (not just unread), identifies messages needing replies, archives handled conversations.
Read-only AWS account hygiene audit — security baseline, unused/orphaned resources, and cost optimization across all configured regions. Produces severity-ranked findings (CRITICAL→LOW) plus a machine-readable findings.json. Cleanup actions are always human-gated, never automatic. Use for cost reviews, security sweeps, recurring account hygiene, or "audit my AWS".
Token-efficient morning briefing. Pre-gathers all data via shell scripts, then presents a unified business dashboard with prioritized actions.
Revenue and costs tracker. AWS spend via aws ce, credits tracker, project revenue stages. Shows burn rate, runway estimate, credits expiring.
SOC 職業分類に基づく
| name | ops-fires |
| description | Production incidents dashboard. Reads ECS health, Sentry errors, CI failures. Offers to dispatch fix agents for active fires. |
| argument-hint | [project-alias|all] |
| allowed-tools | ["Bash","Read","Grep","Glob","Skill","Agent","AskUserQuestion","TeamCreate","SendMessage","TaskCreate","TaskUpdate","Monitor","WebFetch","WebSearch","mcp__sentry__search_issues","mcp__sentry__get_issue_details"] |
| effort | medium |
| maxTurns | 30 |
Before executing, load available context:
Daemon health: Read ${CLAUDE_PLUGIN_DATA_DIR:-$HOME/.claude/plugins/data/ops-ops-marketplace}/daemon-health.json
infra-monitor service status — if not running, pre-gathered infra data may be staleaction_needed is not null → surface it immediately as a potential fireSecrets: AWS credentials are required for ECS/CloudWatch queries.
$AWS_ACCESS_KEY_ID / $AWS_PROFILE env varsdoppler secrets get AWS_ACCESS_KEY_ID --plain (if doppler configured in prefs)password_manager_config.query_cmd from preferences$SENTRY_AUTH_TOKEN → Doppler SENTRY_AUTH_TOKEN → vaultPreferences: Read ${CLAUDE_PLUGIN_DATA_DIR}/preferences.json for secrets_manager config to know which vault to query.
| Command | Usage | Output |
|---|---|---|
aws ecs list-services --cluster <name> --query 'serviceArns' | ECS services | ARN list |
aws ecs describe-services --cluster <name> --services <arn> --query 'services[0].{status:status,running:runningCount,desired:desiredCount}' | Service health | JSON |
aws logs tail /ecs/<service> --since 1h --format short | ECS logs | Log lines (use with Monitor for live) |
| Command | Usage | Output |
|---|---|---|
gh run list --limit 20 --json status,conclusion,name,headBranch,createdAt | Recent CI runs | JSON array |
gh run view <id> --repo <repo> --log-failed | Failed CI logs | Log output |
| Command | Usage | Output |
|---|---|---|
sentry-cli issues list --project <slug> --status unresolved | Unresolved issues | Issue list |
curl -H "Authorization: Bearer $SENTRY_AUTH_TOKEN" "https://sentry.io/api/0/projects/<org>/<proj>/issues/?query=is:unresolved" | API fallback when MCP unavailable | JSON array |
If CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1 is set, use Agent Teams when dispatching multiple fix agents simultaneously. This enables:
Team setup (only when flag is enabled, dispatch phase):
TeamCreate("fire-fixers")
Agent(team_name="fire-fixers", name="fix-[service]", ...)
If the flag is NOT set, use standard parallel subagents.
${CLAUDE_PLUGIN_ROOT}/bin/ops-infra 2>/dev/null || echo '{"clusters":[],"error":"infra check failed"}'
Live anomaly feed from finops-dashboard (spend spikes, idle services,
expired credits, drift detections). High-severity items belong in the
FIRES table alongside infra outages. Falls open to [] if the dashboard
isn't configured.
${CLAUDE_PLUGIN_ROOT}/scripts/finops-bridge.sh anomalies high 2>/dev/null || echo "[]"
${CLAUDE_PLUGIN_ROOT}/bin/ops-ci 2>/dev/null || echo '[]'
${CLAUDE_PLUGIN_ROOT}/bin/ops-external 2>/dev/null || echo '[]'
home_automation is configured in $PREFS_PATH)if jq -e '.home_automation' "${CLAUDE_PLUGIN_DATA_DIR:-$HOME/.claude/plugins/data/ops-ops-marketplace}/preferences.json" >/dev/null 2>&1; then
${CLAUDE_PLUGIN_ROOT}/bin/ops-home snapshot 2>/dev/null || echo '{"configured":true,"error":"home probe failed"}'
else
echo '{"configured":false}'
fi
Analyze the pre-gathered data — including external projects. Then run parallel checks:
gh run list --limit 20 --json status,conclusion,name,headBranch,createdAt 2>/dev/nullauth_expired as HIGH (credential rotation needed), unreachable/degraded as MEDIUM, not_configured as LOW.configured:true) — classify Homey incidents:
/ops:ops-home alarm for details./ops:ops-home status.
If snapshot returned configured:false, skip silently.Classify each issue by severity:
| Severity | Criteria |
|---|---|
| CRITICAL | Service down, DB unreachable, auth broken |
| HIGH | Elevated error rate, deploy stuck, CI main broken |
| MEDIUM | Non-critical service degraded, flaky tests |
| LOW | Warning-level, non-urgent |
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
OPS ► FIRES DASHBOARD — [timestamp]
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
CRITICAL
[service] — [issue] — [since]
HIGH
[service] — [issue] — [since]
MEDIUM
[service] — [issue] — [since]
ECS HEALTH
[cluster] [service] [desired/running] [status]
CI STATUS
[repo] [branch] [workflow] [status] [last run]
SENTRY (top errors, 24h)
[error] [count] [first seen] [project]
EXTERNAL PROJECTS
[alias] [source] [status] [details — e.g. auth_expired, unreachable]
HOME (only if `home_automation` is configured)
[alarm/device] [type — smoke/water/security/offline/energy] [severity] [since]
[If `configured:false`, omit this section entirely]
──────────────────────────────────────────────────────
Use batched AskUserQuestion calls (max 4 options each). Only show relevant actions (e.g., skip dispatch options if no issues found):
AskUserQuestion call 1:
[Dispatch fix agent for [top critical issue]]
[Dispatch fix agent for [second issue]]
[View logs for [service]]
[More...]
AskUserQuestion call 2 (only if "More..."):
[Open Sentry dashboard]
[Open GitHub Actions]
[All clear — nothing to do]
If no fires: show "ALL SYSTEMS OPERATIONAL" with last-checked timestamps.
The pre-gathered CI data is cached and may be minutes-to-hours old. Before
dispatching ANY fix agent, verify the failure is still red on its branch HEAD.
This is defense-in-depth: even after the bin/ops-ci "current-state" filter
(which only emits workflows whose latest run on a tracked branch is failing),
a fix may have landed in the seconds since the cache was written. Dispatching
to a self-resolved fire wastes Sonnet quota — typically 50–150k tokens per
agent before it figures out there's nothing to fix.
For each fire the user selects:
gh run list --repo "$REPO" --workflow "$WORKFLOW" --branch "$BRANCH" --limit 1 \
--json conclusion,databaseId,createdAt --jq '.[0]'
conclusion == "success" → SKIP. Mark task completed with metadata {resolution: "self-resolved-pre-dispatch"}. Do NOT spawn agent.conclusion == "failure" → proceed to dispatch.conclusion == null (in_progress) → wait 30s, recheck once, then proceed if still null.For workflows scoped only to PRs (no main/dev runs), check the PR's combined CI status instead: gh pr checks <num> --repo "$REPO" --json bucket,name.
When user selects to fix an issue, use AskUserQuestion to confirm the scope before dispatching:
Dispatch fix agent for: [issue title]
Severity: [CRITICAL/HIGH/MEDIUM]
Repo: [repo]
Error: [brief description]
The agent will:
- Investigate root cause in [repo]
- Create feature branch with fix
- Open PR for review
[Dispatch agent] [Show me the logs first] [Skip — I'll fix manually]
On confirmation, spawn an Agent with:
Use the agents/infra-monitor.md agent definition for infra issues.
If $ARGUMENTS contains a project alias, filter to that project's services only.
Use Monitor to stream ECS task logs or GitHub Actions runs when investigating fires:
Monitor(command: "aws logs tail /ecs/<service> --follow --since 5m")
Use TaskCreate for each active fire. Update with TaskUpdate as fires are investigated/fixed/escalated.
When diagnosing fires, use WebFetch to check AWS status page (https://health.aws.amazon.com/health/status), Vercel status, or third-party API status pages.
Use WebSearch to find if the error pattern matches a known AWS/infrastructure issue (e.g., "ECS task stopped CannotPullContainerError" → known ECR throttling).
The ops-daemon surfaces two additional fire categories in daemon-health.json:
credential_warnings — tokens/keys expiring within 7 days OR API keys older than 180 days. Fed by offline inspection of preferences.json (*_expires_at, *_created_at fields). No live API calls are made to validate credentials.rate_limit_warnings — integrations currently at ≥80% of their quota window. Fed by counters in rate-limits.json. Resets automatically when the window rolls over./ops:fires lists both alongside Sentry / infra / CI issues. Push notifications are dispatched by the daemon on the first crossing of the threshold — not re-sent until the next day (credentials) or window rollover (rate limits).
CLAIM_KEY: sentry:issue:<short_id> (e.g. sentry:issue:MY-PROJECT-1A2B)
For non-Sentry fires (infra, CI, credential expiry), use:
ci:run:<repo>:<run_id>credential:expiry:<service>CLAIM_KEY="sentry:issue:<short_id>"
ledger query --claim-key "$CLAIM_KEY" --since=-PT24H
If in_progress or done exists, skip the issue. If awaiting_sam exists, surface
it as "fix already staged — needs your decision."
# Claim when beginning to investigate/fix
ledger write \
--claim-key "$CLAIM_KEY" \
--kind "fix" \
--status "in_progress" \
--title "Fire: <issue title>" \
--ttl-sec 7200
# Resolve after fix is applied or escalated
ledger write \
--claim-key "$CLAIM_KEY" \
--kind "fix" \
--status "done" \
--title "Fire: <issue title>" \
--context "fixed: <brief resolution> | escalated: <reason>"