| name | code-review |
| description | Evidence-backed review of source code and machine-consumed implementation artifacts, including tests, configuration, schemas, migrations, build scripts, and Git workspaces, diffs, commits, PRs, or ranges. Use only when the target is code or implementation behavior and the user explicitly asks to find defects or risks, assess quality, verify correctness, requirements, or standards, perform a named review dimension, or produce findings or a review conclusion. Do not use for requests that only read, inspect, explain, map, summarize, or trace how code works or its implementation approach; neutral wording such as "check", "look at", or "检查一下" is not review intent by itself. Do not use to review prose documents themselves. Pair semantic LLM review with the bundled deterministic Runner for frozen inputs, disposition accounting, source-anchor validation, per-Finding challenge coverage, and approval gating; publish only confirmed findings and wait for separate authorization before fixing them. |
Code Review
Use this Skill as the thin orchestration layer for code review. Let Codex understand
intent, contracts, behavior, and risk; let scripts/review.mjs own mechanically
decidable target membership, snapshots, evidence coordinates, declared dispositions,
challenge coverage, input freshness, and conclusion consistency. The Runner does not
dispatch the model and cannot prove that an Agent understood an item or that a
challenge verdict is semantically true or independent; a successful command proves
only the frozen inputs and recorded process state, not semantic completeness or
correctness.
Keep target code, specifications, and Git history read-only until the user separately
authorizes fixes. Runtime and candidate artifacts may be written only outside the
target repository unless the user explicitly chooses another location.
Before invoking the Runner or creating review artifacts, require both an implementation
target and explicit review intent. Requests to find bugs or risks, assess quality,
verify correctness or requirement compliance, check coding or implementation standards,
perform a named review dimension, or produce Findings satisfy the intent gate. Requests
that only ask to read, explain, summarize, map, or trace architecture, data flow,
implementation behavior, or an implementation approach do not. A neutral verb such as
"inspect", "check", "look at", or "检查一下" does not satisfy the gate by itself. For
exploratory requests, do not invoke the Runner or create Finding artifacts; use ordinary
repository exploration or another directly applicable Skill.
The review target must contain at least one software implementation artifact. Treat
prose requirements and design documents only as authority or context for evaluating
that implementation. If the user asks to review a document's own clarity, structure,
consistency, or design quality, this Skill and its Runner are out of scope.
1. Establish authority and target
Read all repository rules governing the target, such as applicable AGENTS.md,
CLAUDE.md, CONTRIBUTING.md, and coding standards. Load only architecture documents
and ADRs relevant to the changed modules, contracts, or call paths. Repository rules
and confirmed requirements override this general baseline.
Resolve the requested target:
- For current changes, freeze
HEAD -> index staged changes and index -> worktree
unstaged changes as independent items, plus untracked files. Never collapse them
into one HEAD -> worktree net diff.
- For a branch, commit, range, or PR comparison, resolve the fixed comparison point
once. A three-dot comparison uses its merge base.
- For a current-state feature review, semantically discover the complete file scope
from entry points, callers, contracts, configuration, and tests before freezing it
as explicit files. This mode does not require Git, a diff, or a baseline.
- Require a Git baseline only for change attribution, regression, omission, or
change-set scope conclusions. Ask when repository evidence cannot identify the
requested target.
Use node <skill-directory>/scripts/review.mjs prepare --repo <repository> for a
workspace, add --base <ref> --head <ref> for a fixed range, or repeat --file <path> for a current-state scope. Read the returned queue_path, not the full
Manifest: the queue contains only the item IDs, paths, changed ranges, metadata, and
exclusions needed for semantic review. Frozen source is available relative to the
queue directory as snapshots/<item_id>.before|after. The Runner retains the full
integrity data outside model context.
Read Deterministic Review Runtime Protocol
only when a command fails, an input is excluded or invalidated, the run must be
resumed or diagnosed, or the trust boundary matters. If Node.js or the Runner is
unavailable, or the target is only remote/pasted and cannot be represented, disclose
UNMANAGED_REVIEW; continue only when useful, never issue APPROVE, and do not
claim complete coverage.
When a diff exceeds roughly 500 lines, summarize and batch the review queue by module or
feature. Group mixed concerns by logical function rather than file order. Every item
must still receive an explicit disposition.
2. Select one primary workflow
Read exactly one primary workflow completely:
| Scenario | Primary workflow |
|---|
| Current workspace, routine PR, commit, range, or named feature review without specification acceptance | Routine review |
| Verify whether current implementation or a defined change set satisfies confirmed specifications, tickets, or acceptance criteria | Acceptance review |
| User explicitly restricts review to security, reliability, architecture, SOLID, performance, correctness, specification compliance, or removal candidates | Focused review |
Default a resolvable general code-review request to routine review. Clarify only when
routine versus acceptance review would materially change the conclusion.
3. Load justified review modules
Load the workflow defaults, then only modules justified by the request or concrete
signals:
| Module | Load when |
|---|
| Correctness and quality | Default for routine; focused correctness, performance, error handling, or edge cases; acceptance changes carrying those risks |
| Security and reliability | Authentication, authorization, input, payments, secrets, writes, transactions, concurrency, external calls, or resource consumption |
| Architecture and standards | Default for acceptance; module or public-contract boundaries, inheritance, new abstractions, large refactors, architecture, or SOLID |
| Specification compliance | Default for acceptance; focused specification, ticket, or acceptance-criteria comparison |
| Removal plan | Deprecated, superseded, disabled, or unused paths; explicit cleanup-candidate review |
Focused reviews stay inside the requested dimensions. Do not load checklist modules
as a substitute for tracing the actual code path.
4. Review semantically and account deterministically
For each pending queue item:
- Inspect the relevant change or current-state file, callers, consumers, contracts,
and tests required by the selected modules.
- Form and challenge concrete defect hypotheses. Distinguish verified behavior from
possibility and omit claims without a reproducible trigger or contract violation.
- Record the item as
reviewed only after that work. This is an Agent declaration,
not mechanical proof that the analysis occurred. Record skipped or failed with
the real reason; never use disposition state to exaggerate semantic coverage.
If a current-state feature scope expands, discard the old run and prepare a new
run with the complete explicit file set. Repository context read only to explain
a changed item need not become a finding target; any file used as a finding location
must belong to the Manifest.
5. Validate candidates, challenge findings, and gate the conclusion
Create one schema-version-2 candidate document conforming to
findings.schema.json. Bind every Finding to its
exact Manifest item_id. Use a line anchor with path, side, range, and
existing_code, or a file anchor containing exact frozen Git metadata changes when
the change has no text hunk. Also include P0-P3 severity, trigger, impact, evidence,
and the smallest safe fix direction.
After all dispositions are current, run validate; correct rejected anchors from
the frozen snapshot or omit the Finding. Any later mark invalidates the validated
set and all challenge decisions, so validate again before finalization. Never relocate
a comment by guesswork. Treat validate as mechanical candidate admission, not proof
that the defect exists.
When at least one candidate validates, read Finding Challenge
Protocol completely. Challenge every validated
candidate exactly once without performing another general review. Use a fresh,
read-only subagent for independent challenges when the protocol's risk routing
justifies the coordination cost and Native delegation is available; otherwise perform
the bounded self-challenge it permits. The verifier may confirm, refute, or declare
insufficient evidence for only the supplied candidate. It must not modify code, emit
new findings, or inherit the full implementation conversation.
Create one challenge document conforming to
challenges.schema.json, then run challenge.
Only the Runner-produced confirmed_findings_path may populate the final Findings
section. Keep refuted candidates in the challenge summary, not the bug list. Keep
insufficient P2/P3 candidates as explicit residual risks. An insufficient P0/P1
candidate or scope_status: expanded blocks APPROVE; do not silently ignore it or
start a recursive full review.
Run finalize with the intended workflow conclusion and use only the Runner-allowed
result. Confirmed P0-P2 findings block approval. Do not translate PARTIAL,
INVALIDATED, excluded inputs, unresolved high-risk candidates, scope expansion, or a
blocked conclusion into approval.
Render the selected workflow's user-facing report from confirmed findings and the
challenge summary. State the frozen scope, workflow and dimensions, candidate counts
by verdict, challenge mode, executed versus merely observed checks, excluded or
incomplete items, unreviewed areas, and residual risks. When no findings confirm, say
so without implying more than the frozen target, declared dispositions, and recorded
challenges prove.
Do not use “the AI can no longer find issues” as a completion test and do not launch a
second comprehensive review merely to seek zero new findings. Stop this review when
the frozen queue is fully dispositioned, every candidate has one challenge decision,
the Runner permits the stated conclusion, required checks have their actual status
reported, and every applicable acceptance criterion in an acceptance review has a
recorded result. Choose APPROVE only when required test, typecheck, lint, build, or CI
checks passed on the current input or were explicitly confirmed not applicable, and
all applicable acceptance criteria passed; use COMMENT when required evidence was
not executed or remains unavailable. APPROVE is bounded to that process state; it is
not proof of global semantic completeness. Then wait for separate authorization
before fixing anything.
Exceptional cases
- Empty scope: report what was checked and remain non-approving; ask about another
reasonable scope when one exists.
- Invalid required baseline: stop comparison review without pass/fail attribution.
- Current-state acceptance without baseline: disclose that history, regression, and
change-set attribution were not reviewed.
- Missing confirmed acceptance source: ask where to find it; omit that axis only
after the user confirms none exists.
- Input drift: treat the run as
INVALIDATED, prepare a fresh run, and do not reuse
prior anchors or coverage.