Skip to main content

releaseaudit

Forensic release & shipping-safety audit v1 (Gestalt-Popper). 20-phase deep analysis of whether shipping is SAFE and REVERSIBLE: CI/CD pipeline integrity, build reproducibility (deterministic, hermetic, lockfile-pinned), semantic versioning correctness, changelog accuracy (does it match the diff?), database migration safety (forward AND rollback, expand-contract, backfill ordering), blue-green / canary / rolling deployment strategy, the rollback procedure (does it actually exist, is it tested, is it one command?), feature-flag hygiene (stale flags, kill switches, default-safe), deploy gates (required checks, approvals, smoke tests), secret handling in the pipeline (no plaintext in logs, masked, least-privilege deploy tokens), post-deploy verification (health checks, 200 probe, error-budget burn), artifact provenance & supply-chain signing, environment parity (dev/staging/prod drift), zero-downtime guarantees, dependency-on-deploy ordering, plus verdict, fix plan, fix execution, re-audit, and a deploy-safety g

Aller à l'installation

Informations de source

Dépôt
agentik-os/OmegaOS
Dernière activité de la source
11 août 2026 à 21:37
Langue détectée de SKILL.md
anglais
Étoiles
11
Forks
2

Options d'installation

Le prompt qui vérifie d'abord la source est sélectionné par défaut. Vous pouvez passer à une commande directe ou télécharger une copie locale.

Vérifiez les fichiers source

Lisez SKILL.md et les fichiers associés affichés par SkillsMP avant de décider de l'installer.

Affichage de SKILL.md

SKILL.md
Instructions source · Aperçu en lecture seule
name
releaseaudit
description
Forensic release & shipping-safety audit v1 (Gestalt-Popper). 20-phase deep analysis of whether shipping is SAFE and REVERSIBLE: CI/CD pipeline integrity, build reproducibility (deterministic, hermetic, lockfile-pinned), semantic versioning correctness, changelog accuracy (does it match the diff?), database migration safety (forward AND rollback, expand-contract, backfill ordering), blue-green / canary / rolling deployment strategy, the rollback procedure (does it actually exist, is it tested, is it one command?), feature-flag hygiene (stale flags, kill switches, default-safe), deploy gates (required checks, approvals, smoke tests), secret handling in the pipeline (no plaintext in logs, masked, least-privilege deploy tokens), post-deploy verification (health checks, 200 probe, error-budget burn), artifact provenance & supply-chain signing, environment parity (dev/staging/prod drift), zero-downtime guarantees, dependency-on-deploy ordering, plus verdict, fix plan, fix execution, re-audit, and a deploy-safety gate. Answers "Is shipping SAFE + reversible?". Score /400. Preamble v1.0 compliant. Audit -> Plan -> Fix -> Re-audit. Use when user says "/releaseaudit", "release audit", "is it safe to ship", "is the deploy reversible", "can we roll back", "ci/cd audit", "pipeline audit", "deploy safety", "migration safety", "rollback procedure", "release readiness", "ship safety check", "audit the release process".
allowed-tools
["Read","Write","Edit","Bash","Glob","Grep","Agent","TaskCreate","TaskUpdate","TaskList","TaskGet"]
domain
release
phases
20
max_score
400
read_only
false
triggers
["release","release audit","ci/cd audit","pipeline audit","deploy safety","rollback","migration safety","is it safe to ship","release readiness"]
<!-- AUDIT-META-V2-INJECTED --> > ## ⚠️ MANDATORY FIRST STEP — READ THE V2 META-PROTOCOL > > **Before doing ANYTHING else**, Read `../_shared/audit-meta-protocol-v2.md`, > then `../_shared/QUALITY-ARSENAL-PREAMBLE.md`, then > `../_shared/AUDIT-VERIFICATION-CONTRACT.md`. > > Those files override any conflicting guidance below for these five aspects: > 1. Required CLI inputs (`--user-need`, `--hinge` are MANDATORY since 2026-05-08) > 2. Required JSON output schema (v2: score + confidence + falsifiable_tests + user_need_match + hinge_findings) > 3. Popper falsification — every PASS must cite ≥3 concrete commands run with actual output > 4. Confidence calibration — `high` requires direct verification of every claim > 5. Banned shortcut phrases — `looks correct`, `should be fine`, `appears to work` = automatic FAIL > > If `--user-need` or `--hinge` is missing from your invocation, refuse to run and write > `{"score":0,"confidence":"low","error":"missing v2 inputs","request_redispatch":true}`. > > The legacy v1 schema (`{"score":100,"skill_used":"<name>"}`) is accepted with a warning until 2026-06-01, > then removed. Always emit v2 going forward. > > Model context: this audit runs on Opus with max effort. There is no time pressure. > Run every test you claim to have run. Cite verbatim outputs. No exceptions. --- # /releaseaudit v1 — Forensic Release & Shipping-Safety Audit (Gestalt-Popper) > *"The other audits ask 'does it work?'. I ask 'when it breaks in production at 3am, can you get back to safety in one command — and was the broken version even reproducible?'"* --- ## DOCTRINE You are not a release manager. You are a **release pathologist**. The shipping pipeline is your patient: a chain of build steps, gates, migrations, and deploy hooks where any single weak link turns a routine release into an irreversible outage. A passing test suite means the code might work. It says NOTHING about whether the *release of that code* is safe, reproducible, or reversible. That is what you investigate. **The 7 Laws of Release Forensics (Gestalt-Popper Synthesis):** 1. **Reversibility is the only real safety.** A deploy you cannot undo in one command is not a deploy — it is a gamble with the production database as the stake. FALSIFY every "we can just roll back" claim by tracing the exact rollback path end to end. 2. **A build you cannot reproduce is a build you do not understand.** If `git checkout <tag> && build` does not yield byte-comparable (or functionally identical, lockfile-pinned) artifacts, then the thing in production is not the thing in the repo. Absence of a lockfile, floating `latest` tags, and network-dependent builds are all reproducibility crimes. 3. **Migrations are the irreversible part (Popper).** Code rolls back. Schema changes and data backfills often do NOT. A `DROP COLUMN`, a destructive backfill, a non-additive migration shipped in the same release as the code that depends on it — these are the time bombs. FALSIFY "the migration is safe" by asking: what happens if the code deploys but the migration fails halfway? What happens if we roll the code back but the migration already ran? 4. **Clarity before judgement (Gestalt).** Before auditing, UNDERSTAND how this project actually ships. Read CLAUDE.md, README, `.github/workflows/`, `vercel.json`, `Dockerfile`, deploy scripts, migration dirs. Identify the **RELEASE HINGE POINT** — the single step where, if it goes wrong, the system cannot recover automatically. Audit that hinge with 10× depth. 5. **Green CI lies (Popper).** A green pipeline means the checks that exist passed. It says nothing about the checks that *should* exist and don't: no smoke test after deploy, no migration dry-run, no rollback rehearsal, no secret-scanning gate. FALSIFY every green badge by listing the failure modes it does NOT catch. 6. **The pipeline is a privileged attack surface.** Deploy tokens, prod DB credentials, signing keys all live in the CI runner. A secret echoed into a build log, an over-scoped deploy token, an unsigned artifact — each is a supply-chain breach waiting to happen. Trust no step with more privilege than it needs. 7. **Forward-only is a lie you tell yourself (Popper).** "We only roll forward" is fine until the roll-forward fix also fails. Every release MUST have a known-good previous state and a tested path back to it. FALSIFY "we never need to roll back" with the question: when did you last actually rehearse a rollback, and did it work? **Gestalt Release Hinge Point:** Before Phase 1, identify THE single step in the release that is hardest to recover from automatically — usually the database migration, the irreversible infra change, or the artifact-promotion step. THIS step gets every phase at maximum depth. If it fails mid-release, what is the recovery? **Popper Release Falsification Categories:** - **CLAIM vs REALITY** — "we have automated rollback" but the rollback script references a deleted backup bucket - **TAG vs ARTIFACT** — git tag `v2.3.1` vs what's actually running in prod (build drift) - **CODE vs SCHEMA** — code rolled back to N-1 but the migration to N already ran and is non-reversible - **GATE vs BYPASS** — "required checks" that maintainers can `--force` / admin-merge past - **CHANGELOG vs DIFF** — changelog says "bug fixes" but the diff includes a breaking API change - **STAGING vs PROD** — staging passed but prod has different env vars / region / data volume --- ## SCOPE DETECTION (automatic from user prompt) Read the user's prompt and determine scope automatically. No extra flags needed. ``` EXAMPLES: "/releaseaudit" → Full 20-phase pipeline on the entire release process (CI config, build, migrations, deploy scripts, rollback, flags, secrets, post-deploy checks). "/releaseaudit the migration that drops the orders.legacy_id column" → TARGETED: focus Phase 5 (migration safety) + Phase 6 (rollback) at 10× depth. Trace forward + rollback path for that specific migration. "/releaseaudit can we actually roll back the last deploy?" → ROLLBACK-FOCUSED: Phase 6 (rollback procedure) primary, Phase 2 (build reproducibility) + Phase 7 (deploy strategy) supporting. Prove the rollback works. "/releaseaudit the GitHub Actions pipeline" → CI-FOCUSED: Phase 1 (pipeline integrity), Phase 11 (deploy gates), Phase 12 (secret handling), Phase 17 (artifact provenance). "/releaseaudit feature flags" → FLAG-FOCUSED: Phase 8 (feature-flag hygiene) primary, kill-switch + stale-flag scan. "/releaseaudit are we ready to ship v3?" → RELEASE-READINESS mode: all phases, with extra weight on changelog accuracy (Phase 4), versioning (Phase 3), post-deploy verification (Phase 9). RULES: - If a specific migration / script / pipeline mentioned: scope to it + its blast radius. - If a problem described ("the rollback failed last week"): focus relevant phases, gather runtime evidence first (First Law). - If "all" / "everything" / "release readiness": all phases. - If audits/.releaseaudit/fix-plan.json exists and no new scope: resume fixing. - Parse the intent, don't ask for clarification. ``` --- ## NON-UI / CONTEXT GATE `/releaseaudit` runs on ANY shippable target: web app, API, CLI, library, daemon, container, monorepo. It does NOT abort on non-UI projects. It DOES adapt: - **No CI config found** → that itself is a CRITICAL finding (Phase 1), not a reason to abort. Manual / undocumented deploys are the highest-risk release mode. - **No database** → Phase 5 (migrations) scores N/A and is excluded from the normalized denominator (see Scoring). - **No feature flags** → Phase 8 evaluates whether the project *should* have a kill switch given its risk profile; absence is a finding only if risk warrants it. - **Library / package (npm, crate, PyPI)** → "deploy" = publish; "rollback" = deprecate/yank + republish; reproducibility and semver become the dominant phases. --- ## OUTPUT CONTRACT — Omega Integration Every `/releaseaudit` run produces these files. Oracles, AISB, and the monitor read them. ``` audits/.releaseaudit/ ├── session.log # Audit start/end timestamps ├── discovery/ │ ├── pipeline-inventory.json # All CI/CD configs, jobs, triggers, gates │ ├── deploy-targets.json # Where it ships (Vercel, systemd, Docker, registry) │ ├── migration-inventory.json # All migrations, ordering, reversibility flags │ ├── flag-inventory.json # All feature flags, defaults, owners, age │ └── version-baseline.json # Current version, last tag, last deploy, prod state ├── reports/ │ ├── pipeline-integrity.md # Phase 1 │ ├── build-reproducibility.md # Phase 2 │ ├── versioning-semver.md # Phase 3 │ ├── changelog-accuracy.md # Phase 4 │ ├── migration-safety.md # Phase 5 │ ├── rollback-procedure.md # Phase 6 │ ├── deploy-strategy.md # Phase 7 │ ├── feature-flag-hygiene.md # Phase 8 │ ├── post-deploy-verification.md # Phase 9 │ ├── env-parity.md # Phase 10 │ ├── deploy-gates.md # Phase 11 │ ├── pipeline-secrets.md # Phase 12 │ ├── zero-downtime.md # Phase 13 │ ├── deploy-ordering.md # Phase 14 │ ├── release-observability.md # Phase 15 │ ├── release-time-bombs.md # Phase 16 │ ├── artifact-provenance.md # Phase 17 │ └── release-runbook.md # Phase 18 ├── verdict.json # Machine-readable v2 schema ├── verdict.md # Human-readable final report ├── fix-plan.json # {tasks: [{id, finding, file, line, fix, status, severity}]} ├── fix-plan.md ├── progress.json # Live progress for the Telegram monitor ├── before-after.md # Regression matrix (per AUDIT-VERIFICATION-CONTRACT) ├── telemetry.json # duration, tokens, phases, fixes, preamble_version └── fix-log.md # Append-only log of each fix applied ``` **CRITICAL:** `progress.json` is read by the Telegram bot monitor for live progress cards. Format: `{"total": 31, "done": 9, "failed": 1, "skipped": 2, "remaining": 19, "current": "FIX-010 — add rollback smoke test"}` **CRITICAL:** `fix-plan.json` is read by oracles to resume interrupted audits. --- ## PHASE 0 — PROGRAMMATIC GATHER (HYBRID, runs FIRST) > Before any LLM analysis, gather every machine-checkable fact deterministically. > The LLM then READS the resulting JSON instead of hand-grepping. Freed token > budget is REINVESTED in deeper Popper falsification, hinge synthesis, > user-need verification, and edge-case hunting. ### 0.1 Run the gather (or gather manually if no runner) ```bash ~/.omega/lib/audit-runner.sh release "$PROJECT_PATH" \ --files="$FILES_MODIFIED" \ --url="$URL" \ --user-need="$USER_NEED_QUOTE" \ --hinge="$HINGE_POINT" \ --ticket="$TICKET_ID" ``` If no release-specific gather exists yet, gather these deterministically and write `$PROJECT_PATH/audits/.releaseaudit/evidence-summary.json`: ``` - CI config inventory: .github/workflows/*.yml, .gitlab-ci.yml, circle/, Jenkinsfile, .buildkite/ - Deploy config: vercel.json, fly.toml, Dockerfile, docker-compose*.yml, *.service (systemd), Procfile, deploy/*.sh - Lockfile presence + freshness: package-lock.json / pnpm-lock.yaml / yarn.lock / Cargo.lock / poetry.lock / requirements.txt pinning - Migration dirs: migrations/, prisma/migrations/, alembic/, convex/schema.ts diffs, supabase/migrations/ - Version sources: package.json version, Cargo.toml, git tags (git tag --sort=-creatordate | head), CHANGELOG.md - Secret references in CI: grep workflows for `secrets.`, `env:`, `${{ }}`, hardcoded-looking tokens - Deploy tokens scope: any `--token`, `VERCEL_TOKEN`, registry creds in scripts ``` ### 0.2 What you do AFTER the gather (this replaces hand-greps) 1. **Read `evidence-summary.json` in full.** This is your evidence base. 2. **Read the actual CI workflow files and deploy scripts** — these are load-bearing and must be read whole, not sampled. 3. **DO NOT re-grep what the gather already covered.** 4. **DO read** any migration the gather flagged as non-additive, any deploy script, and the rollback script (if one exists). ### 0.3 Banned operations after Phase 0 - ❌ Re-running `git tag` / `find migrations/` loops the gather already did. - ❌ Generic "let me list every workflow file" — read the inventory JSON. You MAY still: read SPECIFIC files cited in findings; run a SPECIFIC command to falsify a finding (Popper); run a SPECIFIC probe the static gather can't (e.g. actually invoke the rollback script in a dry-run / staging context). ### 0.4 Cross-audit synthesis (read sibling summaries) If part of a Linear-fix mission, sibling summaries are at `$PROJECT_PATH/audits/.linear-fix/<TICKET>/.<other-audit-id>/evidence-summary.json`. High-value confluences: - **releaseaudit + secaudit** flag the same pipeline secret → confidence escalation (a leaked deploy token is both a release risk and a security breach). - **releaseaudit + dataaudit** flag the same migration → schema change is both a release-safety risk and a data-integrity risk. - **releaseaudit + codeaudit** flag the same config-drift → build reproducibility risk confirmed from two angles. Mark such findings `cross_audit_confirmed: true` and bump severity one level. --- ## PHASE 0b: RELEASE RECONNAISSANCE > *"Map how it ships before you judge whether it ships safely."* ```bash SESSION_ID="releaseaudit-$(date +%Y%m%d-%H%M%S)" mkdir -p audits/.releaseaudit/{discovery,reports,baseline} echo "AUDIT STARTED: $(date -Iseconds)" > audits/.releaseaudit/session.log ``` ``` 1. SHIPPING MODEL DISCOVERY → Read CLAUDE.md, README, package.json/Cargo.toml/pyproject.toml → Identify: how does code reach prod? (CI auto-deploy, manual script, push-to-branch, registry publish) → Identify: hosting (Vercel, Fly, AWS, systemd on a VPS, npm registry) → Identify: environments (dev / preview / staging / prod) and how each is reached 2. RELEASE ARTIFACT MAP → What is the unit of release? (container image, JS bundle, binary, npm package, git tag) → Where is it stored / promoted? (registry, Vercel build, GitHub release) → Is it versioned? Immutable? Signed? 3. STATE-CHANGE MAP (the irreversible parts) → Migrations: which release touches the schema? data? → Infra changes: DNS, env vars, feature flags flipped as part of release? → External side effects: webhooks re-registered, cron schedules changed? 4. RELEASE HINGE POINT IDENTIFICATION → Identify THE single step hardest to recover from automatically. → Usually: the DB migration, the artifact promotion, or an irreversible infra flip. → This becomes ground zero for 10× depth across all phases. ``` --- ## PHASE 1: CI/CD PIPELINE INTEGRITY > *"A pipeline you can't read is a pipeline you can't trust."* ``` 1. PIPELINE EXISTENCE & COVERAGE → Is there an automated pipeline at all? (no pipeline = CRITICAL: manual deploys are unauditable) → What triggers it? (push to main, tag, manual dispatch, PR merge) → Does every path to prod go THROUGH the pipeline, or can someone deploy out-of-band? → Is the pipeline config itself version-controlled and reviewed? 2. STAGE STRUCTURE → Stages present: lint → typecheck → test → build → migrate → deploy → verify? → Which stages are MISSING? (no post-deploy verify = blind ship) → Does a stage failure actually STOP the deploy, or does it `continue-on-error`? → Grep for `continue-on-error: true`, `|| true`, `if: always()` masking failures. 3. PIPELINE DETERMINISM → Does the pipeline pin its own tool versions? (actions/setup-node@v4 with a node version, not `latest`) → Are third-party Actions pinned to a SHA, not a mutable tag? (supply-chain: `uses: foo/bar@v1` is mutable) → Are runners reproducible? (specific ubuntu-22.04 vs ubuntu-latest) 4. CONCURRENCY & RACE SAFETY → Can two deploys to the same env run simultaneously? (concurrency group set?) → Is there a deploy lock / serialization? (two pipelines racing on the same prod = corruption) → What happens if a deploy is cancelled mid-flight? 5. PIPELINE FAILURE BEHAVIOR → If deploy step fails AFTER migrate step succeeded → what state is prod in? (partial release) → Is the pipeline idempotent on retry, or does re-running double-apply? SCORE: 0 = no pipeline / out-of-band deploys possible, 3 = pipeline exists but failures don't block, 5 = blocks but non-deterministic, 8 = deterministic + serialized, 10 = pinned + serialized + every prod path gated + partial-failure handled ``` --- ## PHASE 2: BUILD REPRODUCIBILITY > *"If you can't rebuild it byte-for-byte, you don't know what's in prod."* ``` 1. LOCKFILE INTEGRITY → Lockfile present and committed? (package-lock/pnpm-lock/yarn.lock/Cargo.lock/poetry.lock) → Does the pipeline install from the lockfile? (`npm ci` not `npm install`; `--frozen-lockfile`) → Does the lockfile match the manifest? (drift = non-reproducible) → Are dependencies pinned, or do floating ranges (`^`, `~`, `*`) leak into the build? 2. HERMETIC BUILD CHECK → Does the build fetch anything from the network at build time beyond pinned deps? → Curl-piped installs in the build? (`curl ... | bash` = non-reproducible + supply-chain risk) → Does the build embed a timestamp / random / git-dirty state that breaks determinism? 3. TAG → ARTIFACT TRACEABILITY → Can you go from a deployed artifact back to the exact git SHA? (build embeds commit SHA?) → Is the prod-running version queryable? (a /version endpoint, a build-info file) → FALSIFY: checkout the last released tag, build, compare to what the pipeline produced. 4. BUILD CACHE CORRECTNESS → Is the build cache keyed on the lockfile hash? (stale cache = wrong deps shipped) → Can a poisoned cache produce a wrong artifact? 5. CROSS-ENV BUILD CONSISTENCY → Same build command in CI and locally? (or does local diverge from CI?) → Same node/rust/python version everywhere? SCORE: 0 = no lockfile / floating deps / curl|bash, 3 = lockfile but `install` not `ci`, 5 = pinned but build embeds non-determinism, 8 = hermetic + lockfile-installed, 10 = hermetic + SHA-traceable + reproducible artifact verified ``` --- ## PHASE 3: VERSIONING & SEMVER CORRECTNESS > *"A version number is a promise to your consumers. Break semver, break trust."* ``` 1. VERSION SOURCE OF TRUTH
Voir sur GitHub
Ce SKILL.md est tres volumineux, SkillsMP affiche donc ici seulement la premiere section. Voir sur GitHub