一键导入
flow-next-spec-completion-review
Verify that a spec's completed tasks fully implement the spec requirements. Use at spec completion before close.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Verify that a spec's completed tasks fully implement the spec requirements. Use at spec completion before close.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Drive any UI surface like a real user - a web app, a Chromium-backed desktop app (Electron / WebView2, reached over CDP), or a genuinely native app (macOS AppKit/SwiftUI, or a non-CDP webview) reached via the Cua Driver / Computer Use. Detects the surface, picks the best available driver, degrades gracefully. Use to navigate sites, verify deployed UI, test web or desktop apps, capture baseline screenshots, drive a sign-in flow, scrape data, fill forms, run an e2e check, or inspect current page state. Triggers on "check the page", "verify UI", "test the site", "test this app", "drive the app", "automate this desktop app", "read docs at", "look up API", "visit URL", "browse", "screenshot", "scrape", "e2e test", "login flow", "capture baseline", "see how it looks", "inspect current", "before redesign", "Electron app", "native app".
Carmack-level implementation review of changes via the configured backend. Use when asked to review code or a diff in a flow-next repo.
Open a PR with a cognitive-aid body rendered from flow-next spec state via gh. Use whenever asked to make or open a PR in a flow-next repo.
Single-tick autonomous build-loop conductor. Advances one ready spec one stage per tick, emits PILOT_VERDICT. Use when asked to pilot a spec or backlog.
Carmack-level review of a flow-next spec or plan via the configured backend. Use when asked to review a plan or spec.
Plan a feature into a flow-next spec with tasks in .flow/. Use when asked to plan, spec out, or break down work (fn-N ids).
| name | flow-next-spec-completion-review |
| description | Verify that a spec's completed tasks fully implement the spec requirements. Use at spec completion before close. |
| user-invocable | false |
Workflow is backend-split. Read workflow-common.md for Phase 0 (backend detection + philosophy), then read ONLY the file matching your active backend:
BACKEND=codex → workflow-codex.mdBACKEND=copilot → workflow-copilot.mdBACKEND=cursor → workflow-cursor.mdBACKEND=rp → workflow-rp.mdDo not load the others — only the active backend's file is needed.
Verify that the combined implementation of all tasks in a spec satisfies the spec requirements. This is NOT a code quality review (that's impl-review's job) — this confirms spec compliance only.
Role: Spec Completion Review Coordinator (NOT the reviewer)
Backends (branch on the Phase 0 RP_ELIGIBLE probe):
RP_ELIGIBLE=1: RepoPrompt (rp), Codex CLI (codex), GitHub Copilot CLI (copilot), or Cursor CLI (cursor)RP_ELIGIBLE=0: Codex CLI (codex), GitHub Copilot CLI (copilot), or Cursor CLI (cursor) — rp is macOS-only; never list it in guidance you surface (--review=rp stays accepted)The executable Phase 0 lives in workflow-common.md §"Phase 0: Backend Detection" — Read it and execute it ONCE, before any other bash in this skill. It defines $FLOWCTL (bundled — NOT installed globally; which flowctl fails, expected), probes RP_ELIGIBLE, resolves $BACKEND via the single flowctl review-backend call, and handles the ASK / none cases. Never invoke flowctl review-backend a second time in the same run.
Exception: a --review=<backend> argument (see Backend Selection below) wins — when present, set BACKEND from the flag and skip Phase 0's review-backend call + ASK handling (still run its $FLOWCTL / RP_ELIGIBLE setup lines).
When RP_ELIGIBLE=0 (not macOS, no supported RepoPrompt CLI), never steer the user toward rp: every backend summary, recommendation, or override hint you surface presents only the runnable configured backends codex, copilot, cursor (plus none). Suppression is not a ban: an explicit --review=rp, FLOW_REVIEW_BACKEND=rp, or review.backend=rp still resolves to rp and errors at runtime via require_rp_cli().
Priority (first match wins):
--review=rp|codex|copilot|cursor|none argumentFLOW_REVIEW_BACKEND env var — bare backend (rp, codex, copilot, cursor, none) OR spec form (codex:gpt-5.4:xhigh, copilot:claude-opus-4.5, cursor:gpt-5.5-high).flow/config.json → review.backend (same bare / spec forms)Check $ARGUMENTS for:
--review=rp or --review rp → use rp--review=codex or --review codex → use codex--review=copilot or --review copilot → use copilot--review=cursor or --review cursor → use cursor--review=none or --review none → skip reviewIf found, use that backend and skip all other detection.
No --review flag → $BACKEND comes from workflow-common.md Phase 0 (executed once per the Preamble): the single flowctl review-backend "$SPEC_ID" call with ASK handling included. Do not re-resolve here.
When RP_ELIGIBLE=0, omit the rp line below from any guidance you surface (explicit --review=rp still honored):
gpt-5.5). FLOW_CODEX_MODEL / FLOW_CODEX_EFFORT env vars, or --spec codex:gpt-5.4:xhigh.FLOW_COPILOT_MODEL / FLOW_COPILOT_EFFORT env vars, or --spec copilot:claude-opus-4.5:xhigh.cursor-agent, cross-platform); reaches gpt-5.5-high (1M-ctx default), the gpt-5.3-codex family, composer-2.5, and claude-opus-4-8-thinking-high via a Cursor subscription. FLOW_CURSOR_MODEL env var, or --spec cursor:gpt-5.5-high. Cursor folds reasoning effort into the model name — no effort field.Spec grammar: backend[:model[:effort]] — FLOW_REVIEW_BACKEND and .flow/config.json review.backend both accept this. Examples: codex, codex:gpt-5.2, copilot:claude-opus-4.5:xhigh, cursor:gpt-5.5-high (cursor takes model only — no :effort). Per-spec default_review (set via flowctl spec set-backend) overrides env.
For rp backend:
setup-review (5-15 min, DO NOT RETRY) - handles window selection + builder atomically--new-chat after first reviewFor codex backend:
$FLOWCTL codex completion-review exclusively--receipt for session continuity on re-reviewsFor copilot backend:
$FLOWCTL copilot completion-review exclusively--receipt for session continuity on re-reviews (session only resumes when prior receipt has mode == "copilot")--spec backend:model:effort flag, per-spec default_review, FLOW_REVIEW_BACKEND spec, FLOW_COPILOT_MODEL / FLOW_COPILOT_EFFORT env vars, registry defaultsFor cursor backend:
$FLOWCTL cursor completion-review exclusively--receipt for session continuity on re-reviews (session only resumes when prior receipt has mode == "cursor")--spec cursor:<model> flag, per-spec default_review, FLOW_REVIEW_BACKEND spec, FLOW_CURSOR_MODEL env var, registry default (gpt-5.5-high). No effort — Cursor bakes effort into the model name; cursor:<model>:<effort> is rejectedFor all backends:
REVIEW_RECEIPT_PATH set: write receipt after SHIP verdict (RP writes manually after fix loop; codex writes automatically via --receipt)<promise>RETRY</promise> and stopFORBIDDEN:
Arguments: $ARGUMENTS
Format: <spec-id> [--review=rp|codex|copilot|cursor|none]
fn-1 or fn-22-53k--review - Optional backend overrideREPO_ROOT="$(git rev-parse --show-toplevel 2>/dev/null || pwd)"
Parse $ARGUMENTS for:
fn-* → SPEC_ID--review=<backend> → backend override$BACKEND was already resolved by workflow-common.md Phase 0 (Preamble) — do NOT re-run it.$BACKEND | File to read |
|---|---|
codex | workflow-codex.md |
copilot | workflow-copilot.md |
cursor | workflow-cursor.md |
rp | workflow-rp.md |
Do not read the other backend files. Each is self-contained for its backend; loading the others wastes context.
Follow the phases in the per-backend file end-to-end. Each file owns its own Identify → Execute → Verdict → Receipt steps (and, for RP, the full Phase 1-4 setup-review (5-15 min, DO NOT RETRY) / chat-send (2-10 min, DO NOT RETRY) / receipt build).
CRITICAL: Do NOT ask user for confirmation. Automatically fix ALL valid issues and re-review — our goal is complete spec compliance. Never use the plain-text numbered prompt in this loop.
MAX ITERATIONS (backend-agnostic — applies to ALL backends: rp, codex, copilot, cursor): keep an iteration counter in agent context, starting at 0. Each fix+re-review cycle increments it. When the counter reaches ${MAX_REVIEW_ITERATIONS:-4} (default 4; env-overridable, configurable in Ralph's config.env) and the verdict is still NEEDS_WORK, BREAK the loop and escalate: surface the surviving gaps to the caller and stop (in Ralph mode output <promise>RETRY</promise> so the next iteration starts fresh). Never loop unbounded. The per-backend workflow files defer to this cap.
If verdict is NEEDS_WORK, loop internally until SHIP or the iteration cap:
git add --all)flowctl codex completion-review (receipt enables context)flowctl copilot completion-review (receipt enables context; must be mode == "copilot" to resume)flowctl cursor completion-review (receipt enables context; must be mode == "cursor" to resume)$FLOWCTL rp chat-send (2-10 min, DO NOT RETRY) --window "$W" --tab "$T" --message-file <literal re-review path from workflow-rp.md's fix loop> (NO --new-chat; stdout redirected to the same literal response file, Read once)<verdict>SHIP</verdict> — or the MAX ITERATIONS cap above breaks the loop (escalate with surviving gaps)CRITICAL: For RP, re-reviews must stay in the SAME chat so reviewer has context. Only use --new-chat on the FIRST review.
flowctl <backend> completion-review self-writes completion_review_status / completion_reviewed_at from the parsed verdict on codex/copilot/cursor (fn-112). Without a write somewhere, a standalone completion review leaves completion_review_status: unknown, which keeps flowctl ready --require-completion-review demanding a review (pilot's gate), feeds make-pr's Open-items / draft heuristic stale state, and blocks tracker-sync's terminal verified rung. The standalone command remains for rp and for repairing a missed write:
# Final verdict resolved to SHIP → ship; NEEDS_WORK at the iteration cap → needs_work.
# Skip when the backend handler already wrote status (codex/copilot/cursor).
$FLOWCTL spec set-completion-review-status "$SPEC_ID" --status ship --json # on SHIP
$FLOWCTL spec set-completion-review-status "$SPEC_ID" --status needs_work --json # on NEEDS_WORK at cap
On rp (or if the JSON payload lacks plan_review_status/completion_review_status), write on BOTH terminal paths (SHIP and capped-NEEDS_WORK).