| name | gcp-triage |
| description | Monitor three.ws production on Google Cloud Run and fix what the sweep finds. Use when the user asks to check production, diagnose an outage or error, "what's wrong with three.ws", read production logs, or run the monitoring loop. Broad questions get the deep sweep (healthz + all-service logs + version, TLS, fleet readiness, live pages, crons, DB migrations, wallets), then the fixes each class allows. |
| when_to_use | Production health checks, log diagnosis, error triage, and the recurring monitor-and-fix loop. For a raw log view only, `npm run logs` is enough without this skill. |
| license | MIT |
| metadata | {"category":"ops/production","cross-platform-safe":false,"pack":"three-ws-ops"} |
GCP production triage: monitor, classify, fix
The vercel logs era is over; production is 20 Cloud Run services in project
aerial-vehicle-466722-p5 (region us-central1). This skill is the loop an
agent runs to answer "is production healthy?" and to act on what it finds.
gcloud is already authenticated in this workspace.
Step 1: run the monitor
When the user asks "what's wrong with three.ws?" (or anything equally broad),
run the DEEP sweep; it checks everything, not just what happened to log:
npm run triage:gcp -- --json --deep --since 6h
For a quick pulse check mid-incident, the fast form is still fine:
npm run triage:gcp -- --json --since 1h
The base monitor merges three signals and classifies every distinct problem
signature:
https://three.ws/api/healthz: the platform's own subsystem roll-up
(database, cache, Helius, x402 ring, world, sniper).
- A WARNING+ log sweep across every Cloud Run service, fingerprinted so
repeats group into one finding.
- HTTP request logs: 5xx groups per route, 429s.
--deep adds nine concurrent read-only probes, each wrapping an existing
standalone audit, normalized into the same findings stream:
| probe | what it proves | wraps |
|---|
version | /api/version answers; deployed commit is in local history; deploy lag counted | built-in |
tls | certs on three.ws + world.three.ws have >21 days left | built-in |
fleet | every Cloud Run service's latest revision is Ready | gcloud run services list |
pages | every page advertised in data/pages.json serves on the live site | scripts/check-pages.mjs --base |
cron-drift | vercel.json crons match live Cloud Scheduler jobs | scripts/check-cron-drift.mjs |
cron-liveness | every cron routes, resolves a handler, and loads | scripts/audit-cron-liveness.mjs --static |
db-migrations | no migration is pending against the database | scripts/apply-migrations.mjs --check |
service-wallets | signer wallets above SOL floors; advertised x402 keys match secrets | scripts/audit-service-wallets.mjs |
custodial-keys | no funded custodial wallet sits behind an undecryptable key | scripts/audit-custodial-key-health.mjs |
A probe that cannot run becomes an investigate finding (a blind spot is not
"healthy"); one skipped for missing local secrets is reported as skipped.
--skip pages,custodial-keys drops probes when you must (e.g. re-running in a
tight loop). The deep sweep takes a few minutes; that is the point.
Exit 0 = healthy or self-healing noise only. Exit 1 = findings[] contains
actionable items. Each finding carries class, count, services,
sample, and a concrete action. Deep findings carry
signature: "deep-<probe>".
Surfaces even the deep sweep does not cover, worth running when the complaint
points at them: npm run smoke:mcp (remote MCP endpoints),
npm run smoke:x402-facilitator (facilitator verifies a real signed payment
without broadcasting it, so it is safe to run any time), npm run audit:web
or audit:web:login (real-browser page audit with the QA account).
Step 2: act per class, in this order
| class | what it means | what you do |
|---|
owner | money / billing / security credential | Do NOT act. Put the exact command from docs/ops/production-log-triage.md in your report. |
env-action | fix is a config-only Cloud Run env/resource change | Apply it now. Config-only gcloud run services update is pre-approved (CLAUDE.md). ALWAYS --update-env-vars (merges); NEVER --set-env-vars (replaces the whole set). |
investigate | unknown signature or 5xx group | Root-cause it (Step 3). Fix the code, add tests, commit locally. Deploys need owner approval, so prepare the one-command ship and say so in the report. |
self-healing | documented graceful degradation | No action. Only escalate if the same finding persists across several runs (compare firstSeen, it should not span many hours). |
Step 3: root-causing an investigate finding
Step 4: close the loop
- Fixed something in code? Commit locally with explicit paths (never
git add -A). Do not push or deploy without owner approval; leave the ship
as one command in the report.
- Found a NEW signature that is expected degradation (a fallback firing
correctly)? Add it to
KNOWN_SIGNATURES in
scripts/gcp-triage.mjs AND document it in
docs/ops/production-log-triage.md,
so the monitor never flags it as investigate again. That is how this
system learns.
- Healthz unreachable? That is an outage: check
three-ws-api revisions
(gcloud run revisions list --service three-ws-api --region us-central1)
and the load balancer per the production runbook before anything else.
Report format
Lead with the verdict (healthy / degraded / outage). Then: what you fixed,
what you committed, what needs the owner (exact commands), and what is
self-healing noise. Solana-facing findings first.