Skip to main content

gate

Pre-push quality gate. Runs parallel review agents over a changeset (reuse, correctness, quality, i18n, wiring, regression, tests, UX, performance) plus an optional external second-opinion review, then fixes what they find. Use before pushing code or creating PRs.

Informações da origem

Repositório
tphakala/claude-gate-skill
Última atividade na origem
21 de agosto de 2026 às 11:35
Idioma detectado do SKILL.md
inglês
Estrelas
1
Forks
0

Opções de instalação

Por padrão, está selecionado o prompt que primeiro revisa a origem. Você pode mudar para um comando direto ou baixar uma cópia local.

Revise os arquivos de origem

Leia o SKILL.md e os arquivos complementares exibidos pelo SkillsMP antes de decidir se vai instalar.

Exibindo SKILL.md

SKILL.md
Instruções da origem · Visualização somente leitura
name
gate
description
Pre-push quality gate. Runs parallel review agents over a changeset (reuse, correctness, quality, i18n, wiring, regression, tests, UX, performance) plus an optional external second-opinion review, then fixes what they find. Use before pushing code or creating PRs.
# Gate: Pre-Push Quality Gate Parallel code review over a changeset, then fixes. This file is the controller's runbook. The checks themselves live in `reference/patterns-agent<N>.md`, one file per agent, each holding that agent's checklist and its worked examples. Do not restate checks here. ## Phase 0: Mechanical floor (deterministic, before any agent) Run these first. They are cheap, and each exists because a check that already existed was skipped or run wrong. **Read `reference/tooling.md`** for what each decides and what it leaves to an agent. ```bash bash {SKILL_DIR}/scripts/discover.sh > /tmp/gate-facts.json # this repo's shape, with evidence export GATE_DISCOVER_JSON=/tmp/gate-facts.json # so every script discovers once bash {SKILL_DIR}/scripts/verify.sh # the project's OWN checks, with CI's flags bash {SKILL_DIR}/scripts/static-checks.sh # matcher-decidable checks over the diff bash {SKILL_DIR}/scripts/contract-diff.sh origin/main # Agent 6 surfaces 1-4, by set arithmetic ``` `discover.sh` replaces every hardcoded layout assumption: where the frontend roots are, where the translations live, which CSS framework and component libraries are declared, which generated artifacts have drift guards, whether config reloads at runtime, and which issue tracker to file against. **Every fact carries the evidence that produced it, or an explicit null and the reason.** There is no guessed default anywhere, because a wrong guess and a correct find are indistinguishable downstream. A null is a real answer and it means no consumer may suppress anything. `verify.sh` reports PASS / FAIL / **SKIP**. A SKIP is not a pass: list every skipped check and its reason in the Phase 3 summary, and never call a changeset mechanically green while a check covering it did not run. `static-checks.sh` splits findings into **decided** (the construct IS the defect) and **candidate** (triage still needed). Hand both to the agents as ALREADY-FOUND; both still go through Step 3.3. ## Phase 1: Scope & Plan ### Step 1.1: Build the scope manifest (deterministic) ```bash bash {SKILL_DIR}/scripts/scope.sh > /tmp/gate-scope.json # or pass a base ref cat /tmp/gate-scope.json ``` The script decides which files are reviewable, classifies them by area (by CONTENT, not by directory name), sets the size class, and packs bundles, all in code, so no file is silently skipped. That is the coverage guarantee, and it is deliberately not a model judgment. **It resolves the review range and PRINTS which one it chose to stderr.** Staged changes, else unstaged, else this branch's commits since its merge-base. That last case is the normal pre-push one, and every script used to resolve it as an empty range: `verify.sh` then tested nothing and every trigger keyed on an added diff line was suppressed. Read the line it prints; if the range is wrong, pass the base explicitly. Exit 3 means there is genuinely nothing to review. Manifest fields: - `totals`: `files`, `added`, `deleted`, `reviewable`, `reviewable_lines`, `tests`, `churn_lines`. - `size_class`: `small` or `large`. Drives fan-out. Computed from `churn_lines` (rename-aware) and reviewable counts, so a moved file, lockfiles, generated and vendored code cannot inflate it. `renamed_files` flags moves: keep both halves of a move in ONE shard, or each is blind to the other. - `files[]`: `path`, `area`, `kind` (A/M/D), churn, `reviewable`, `is_test`, `has_tests`. - `routing`: area to reviewable paths (`go`, `rust`, `frontend`, `i18n`, `config`, `db`, `api`, `other`), plus cross-cutting `perf`: changed non-test source a benchmark for its language covers (Go `func Benchmark` or a `.s` file; TS/JS/Svelte a Vitest `*.bench.*`; Rust a crate-level `benches/`). Every changed hand-written `.s`/`.S` file is ALWAYS in `perf`, benchmark or not: assembly exists only for performance. - `bundles[]`: review units for fan-out, one area each, within size caps (~8 files / 500 changed lines; an oversized single file is its own bundle, flagged `oversized`). Paths are repo-relative. Get the root with `git rev-parse --show-toplevel` and prefix it when handing absolute paths to agents. Also capture the raw diff for the agents that need it. **Fallback:** if `totals.reviewable == 0`, review the files the user named or that you edited earlier in this conversation, treat the changeset as small, and skip the bundling logic. ### Step 1.2: Extract the change summary (what's new) From the diff, extract the structured what's-new list, keyed by file. This is the shared context every cross-file agent starts from: - new struct/type declarations and new fields on existing structs - new functions/methods and new exported symbols - new event emission calls - new options/flags on config/settings structs and new config keys - new API request/response fields (json tags) and new endpoints - new or changed DB models / migrations The manifest plus this list is the **shared change summary** passed to every agent in Phase 2. ### Step 1.3: Match the escaped-defects ledger (deterministic) ```bash bash {SKILL_DIR}/scripts/ledger-match.sh /tmp/gate-scope.json "<what's-new keywords>" ``` The script matches the manifest's areas and file extensions plus your keywords against each ledger entry's applies-when trigger and prints only the matching entries. Fold each match into the relevant agent's prompt for this run, so past misses become present coverage. Do not skim the ledger by hand; it is 136 entries and the script is the triage surface. **Optional accelerator: a cross-project memory bank.** If you have a memory/recall tool wired up (any persistent store you can query), recall cross-project lessons matching this changeset to rank the most relevant few. **Read `reference/tooling.md` section "Memory bank"** for the suggested query shape and tags. This step is entirely optional; skip it silently if you have no such tool. Memory NEVER decides which files are reviewed and never relaxes the coverage contract; it only prioritizes lenses. ### Step 1.4: Risk plan (large changesets only) If `size_class == "large"`, dispatch one planner agent with the full diff and manifest. Strict JSON: ```json { "intent": "1-2 sentences on what this change does", "risk_hotspots": [ {"file": "path", "lines": "40-72", "why": "...", "severity": "high|medium|low", "lens": "correctness|reuse|quality|integration|regression"} ], "cross_file_watch": ["short notes on wiring or multi-site concerns to verify"] } ``` `risk_hotspots` is ranked high to low; `lens` routes each hotspot to an agent; `cross_file_watch` seeds Agents 5 and 6. For `small`, skip the planner. ## Phase 2: Launch Review Agents ### NEVER pass `name` to the Agent tool **Dispatch review agents WITHOUT the `name` parameter. Passing it destroys the entire fan-out, silently.** Root-caused 2026-07-20 after three consecutive runs in which every dispatched agent returned nothing. `name` is not a label. It switches the agent to `taskKind: "in_process_teammate"`, whose plain-text output goes to its transcript and nowhere else, so a named agent told to report findings as text does exactly that and the text is discarded. Intended harness behavior, not a bug to work around, and messaging the agents afterwards does not recover it. If a run produces `idle_notification` with no body, check for a stray `name` first. Case history in `reference/tooling.md`. ### Step 2.0: Slice each agent's checklist (deterministic) ```bash # Tier B and scoped agents: slice against the whole changeset. bash {SKILL_DIR}/scripts/triggers.sh --scope /tmp/gate-scope.json --out /tmp/gate-checklists # Tier A shards: slice PER BUNDLE, once per bundle, and hand each shard its own file. bash {SKILL_DIR}/scripts/triggers.sh --files "<bundle's files, comma-separated>" \ --agents 1,2,3 --out /tmp/gate-checklists/<bundle-id> ``` Each agent then reads `/tmp/gate-checklists/agent<N>.md` INSTEAD of its patterns file. Slice per bundle: against the whole changeset each shard gets every area's checklist, which is barely a slice. Suppression is fail-open. An unclassified item is always live, the security and data-loss classes are never gated, and any error yields the unsliced file. Each slice states what it suppressed and why. **Tell every agent to report a suppressed item that looks like it should have applied**; that is a trigger-table bug and it must be visible rather than worked around. ### The two tiers - **Tier A, file-local (Agents 1, 2, 3):** reason within a file. Sharded across bundles when large. - **Tier B, cross-file (Agents 5, 6):** reason across the whole change. **Never sharded.** Isolating them to a bundle blinds them to multi-site fixes, sibling inconsistency and wiring gaps, which are the gate's highest-value findings. ### Agent dispatch table Every agent reads its own `reference/patterns-agent<N>.md`, which carries its checklist and examples. Tell each agent its areas from `routing` and let it select the matching sections. | # | Agent | Tier | Dispatch when | Input | Patterns file | |---|---|---|---|---|---| | 1 | Reuse & Efficiency | A | always | full diff + reviewable file list | `patterns-agent1.md` | | 2 | Correctness & Safety | A | always | full diff + files | `patterns-agent2.md` AND `patterns-agent2-part2.md` (both) | | 3 | Quality & Patterns | A | always | full diff + files | `patterns-agent3.md` AND `patterns-agent3-part2.md` (both) | | 4 | i18n Translation Integrity | scoped | `routing.i18n` or frontend changed | translation files; runs `scripts/i18n-check.sh` first | `patterns-agent4.md` | | 5 | Integration & Wiring | B | always | what's-new + `cross_file_watch` + manifest + paths. **No raw diff**; it reads full files | `patterns-agent5.md` | | 6 | Regression & Backward Compat | B | always | full diff AND paths AND what's-new | `patterns-agent6.md` | | 7 | Test Quality | A | `totals.tests > 0` | `is_test`/`has_tests` files + hunks | `patterns-agent7.md` | | 8 | UX/UI & Accessibility | scoped | `routing.frontend` or `routing.i18n` non-empty | frontend files in full, plus the full changed-frontend path list | `patterns-agent8.md` | | 9 | Performance | A | `routing.perf` non-empty, or you judge the change perf-sensitive | perf-flagged files + hunks; one shard per language | `patterns-agent9{,-rust,-ts}.md`, by language | Agent 8 is **report-only**; it never edits (Phase 3 handles its findings). Agent 9 labels each finding `mechanical-fix` or `report-only`. Agents 8 and 9 both emit COVERAGE tables. **Agent 9 routing:** launch one shard per language, each reading only its own file. Several Go rules INVERT in JS and several have no Rust analogue at all; each file says which. **When Agent 9 runs, Agent 1 defers micro-performance to it** and keeps semantic waste (N+1 fetches, missed parallelism, fetching for hidden UI). ### Dispatch by size class **`small` (or fallback):** one agent per dimension, concurrently in a single message. Tier A gets the full diff plus the reviewable file list. Do NOT fold or merge dimensions to save agents, even on a tiny changeset. Each dimension is a distinct lens, and the reuse lens in particular catches misses a quality-focused pass glosses over (a redundant install of an already-declared dependency, a duplicated helper). **One agent per dimension is the floor, not a target to optimize below.** **`large`:** fan out Tier A. For each entry in `bundles[]`, launch one Agent 1, one Agent 2 and one Agent 3 scoped to that bundle. Give each shard its bundle's files as absolute paths, the diff hunks for those files, the shared change summary, the bundle `area`, and the `risk_hotspots` whose `file` is in its bundle. Mark `oversized` bundles so that shard reads the file in full. Launch Tier B once each over the full manifest, never sharded. Shard Agents 7, 8 and 9 only on very large changesets, and **always pass every Agent 8 shard the full changed-frontend path list** so its cross-component consistency pass is not blinded by sharding. Launch as many as fit in one message; the harness queues the rest. ### The coverage contract **Per-file accounting.** Every Tier A agent, sharded or not, must end its report with a COVERAGE table: one row per assigned file, status `FLAGGED (n)`, `CLEAN`, or `N/A: <reason>`. Every assigned file must appear. A skipped file then surfaces as a gap in Phase 3 instead of vanishing. **Checklist-loaded line.** Every agent must open its report with the patterns file it read and how many checklist items it loaded. The checklist lives in that file, not in the prompt, so a failed Read means the agent reviewed nothing. A missing line, or a count of zero, means NOT DELIVERED: re-dispatch it. **Delivery check (mandatory, before Phase 3).** A report counts as RECEIVED only when its text, with its COVERAGE table, is in hand as a tool result. Say "COVERAGE table" to each agent in those words and state that a findings-summary table is NOT a substitute: when the 2026-07-20 reports were recovered, 2 of 5 had produced complete reviews but substituted their own table format, so no file was ever listed as reviewed-and-clean and the reconciliation had nothing to consume. Count the reports actually received. If that number is below the number dispatched, say so in the summary in those words, treat the missing agents' files as UNREVIEWED, and dispatch the catch-up pass. ### Prompt-authoring rules **Do not transfer your own blind spot.** Example inputs, file lists and the what's-new summary you hand an agent are STARTING POINTS, not the coverage boundary. Never let a curated example list become the agent's whole search space. State the *semantics* of what changed ("the old code used `url.Parse`; derive its full accepted-input domain and test against it") and tell the agent explicitly to expand beyond any examples you gave. A hand-authored input list silently caps the agent at the author's blind spot, which is exactly how the `ws://host/path` regression slipped multiple passes whose prompts all enumerated the same path-less corpus. **Restrict repo-mutating git commands.** "Do not edit any file" does NOT cover `git stash`, `checkout`, `reset`, `restore` or `clean`, which mutate repo state with no file edit. Say "do not run any git command that changes HEAD, the index, or the working tree; compare versions in a scratch copy outside the repo", and give `git archive <ref> | tar -C /tmp/x -x` as the safe recipe. Enforcement is a post-dispatch `git status --porcelain` check reading BOTH columns, not the prompt. **Every agent classifies each finding:** - **[CHANGED]**: in lines added or modified by this diff. - **[PRE-EXISTING]**: in surrounding untouched code, found while reviewing context. Agents review the full context of changed files, not only diff lines. ### Visual-render gate for layout-bearing frontend changes No agent can see a broken layout. Every agent reviews code statically and jsdom asserts attributes, not computed layout, so a purely visual regression (a grid child missing `col-span-*` and collapsing to one track, a broken flex row, a duplicated title, overflow) is invisible to the entire gate. That is how a prior run shipped analytics pages squished into a 1/12 column. Flag the changeset now if Step 3.6's trigger fires; **the gate then cannot pass on static green alone**. ### Background: external second-opinion review (optional, pluggable) If you can dispatch the changeset to an external reviewer, a capable non-Claude model reached over a CLI or an MCP tool, launch it in the background at dispatch time and collect it in Phase 3, and never report while it is unreconciled. This step is optional: if you have no such reviewer wired up, skip it and rely on the in-process agents. **Read `reference/tooling.md` section "Background: external second-opinion review"** for the prompt, a CLI-fallback shape, the concurrency constraint that stops a background reviewer colliding with Agent 6's own cross-check, and why an empirical, suite-running model and a fast-reasoning model catch disjoint issue classes on high-risk diffs. **Its findings are LEADS, not findings.** Reproduce each claim against the actual code before acting; in a documented case all three of an external reviewer's contested claims were false. The tell is a hedge like "likely calls X" where the call does not exist. Applies to external review bots in a push-review cycle too. ## Phase 3: Aggregate, Fix & Report Wait for all agents, and collect the background external-review job if one was launched. Never report while it is unreconciled. ### Step 3.1: Verify coverage Collect the COVERAGE tables from all Tier A agents (1, 2, 3, plus 7 and 9 when they ran). Take the union of files marked FLAGGED / CLEAN / N/A and compare against `files[]` where `reviewable == true`. For any reviewable file with no status from any Tier A agent, dispatch a catch-up agent (correctness + quality) over just those files before continuing. Record "X/Y reviewable files reviewed". Agent 8's COVERAGE table confirms the UX pass was not skipped but does not participate in the catch-up, since Tier A already covers those files for fixes. ### Step 3.2: Filter false positives Apply the static list at the end of this file first. Then, if you have an optional memory bank (Step 1.3), do ONE recall of this project's conventions and confirmed non-issues, for example: ```text recall: "intentional patterns, conventions, and confirmed non-issues a reviewer should NOT flag in this project; things explained as deliberate", tagged to this project. ``` Suppress or downgrade a finding ONLY when a returned memory tagged as a convention or a confirmed false-positive clearly and SPECIFICALLY covers it (same construct, same rationale). Vague topical overlap is not enough; a memory confirmed more often is stronger evidence. Hard guardrails: memory may only suppress or downgrade, **never create** a finding (that is Phase 1's job) and **never suppress a Critical/High security, data-loss or panic finding**, whatever the convention says. Every memory-driven suppression is listed in the summary with the matching learning quoted and its source and confidence cited, so over-suppression is auditable, never silent. ### Step 3.3: Triage **Triage on the finding's ARGUMENT, not the agent's severity LABEL.** Severity is the agent's opinion; the body is the evidence. Agents routinely self-limit a correct finding with a defensible hedge ("consistency call, not a bug", "test-only", "the comment is accurate"), and that label then does the triage instead of you. Escalate to at least Medium and FIX, whatever the label says, when the body shows the changeset treats the SAME hazard two ways: eliminating it structurally in one file while resolving it in prose in another. "Test-only" is not a defence. **A finding you defer is one you have chosen to have a reviewer find.** Deferring is only correct when you would also be content for no reviewer to find it. Weight "constructor accepts a config that cannot work" at Medium or above: the failure surfaces far from its cause. Defer against the CHEAPEST fix, not the first (`reference/tooling.md`, "The deferral test"). Four categories: **a) In-scope (introduced or touched by this diff).** Fix it, at the invasiveness the RE-DERIVED severity earns, per the ladder below. For Critical/High, first re-read the cited code and confirm the failure is real in the actual source: agent reports can be plausible-but-wrong, and a fix applied to a misread finding is itself a regression. If it does not reproduce by reading, do not fix; report it as unconfirmed. **Match the fix's invasiveness to the severity: 23% of the ledger is fix-wave damage.** | Re-derived severity | Fix you are allowed to make | |---|---| | Critical / High | Whatever correctness genuinely requires, including a structural change. | | Medium | A minimal local edit. No new abstraction, no rerouted call sites, no behaviour change. |
Ver no GitHub
Este SKILL.md e muito grande, entao o SkillsMP mostra aqui apenas a primeira secao. Ver no GitHub