원클릭으로
goat-critique
Use when a decision or analysis needs multi-lens critique to surface blind spots before shipping.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
메뉴
Use when a decision or analysis needs multi-lens critique to surface blind spots before shipping.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
SOC 직업 분류 기준
Use when evaluating test coverage gaps, planning test strategy, or assessing testing risk for code changes.
Use when reviewing a diff, PR, or set of code changes, or auditing a codebase area for quality issues. Triggers: 'review this', 'code review', 'audit X', 'look at these changes'.
Use when assessing security implications of code changes, architecture decisions, or new features.
Use when starting a non-trivial implementation that needs structured task breakdown with progress tracking.
Use when you describe an outcome and need the right goat-* workflow chosen for you.
Use when diagnosing a bug, unexpected behaviour, system failure, or unfamiliar code that needs structured investigation.
| name | goat-critique |
| description | Use when a decision or analysis needs multi-lens critique to surface blind spots before shipping. |
| goat-flow-skill-version | 1.14.0 |
Read .goat-flow/skill-docs/skill-preamble.md and .goat-flow/skill-docs/skill-conventions.md for shared conventions before proceeding.
Use when a concrete artifact deserves multi-perspective critique before shipping: plan, security assessment, debug hypotheses, review findings, test strategy, architecture proposal, or refactor approach.
/goat-review for a trivial artifact; explicit /goat-critique still runs in full.| Excuse | Reality |
|---|---|
| "The artifact is trivial - a quick critique would cover it" | Quick mode was tried and removed. A single reviewer running lens passes in one context is self-talk under three labels, not multi-perspective critique. |
| "All three agents agree so it must be right" | Consensus without orchestrator verification is unverified self-declaration. The orchestrator's job is to verify claims, not count votes. |
| "Inline role-play is faster than spawning agents" | Agents that role-play SKEPTIC/ANALYST/STRATEGIST inline produce indistinguishable perspectives. Isolated context is what makes findings independent. |
| "Closing checks happen after the main answer - skip them" | End-of-task rules have near-zero voluntary compliance. Phase 5.5 meta-audit and outcome capture exist because post-deliverable steps get skipped. |
Direct invocation is binding. $goat-critique or /goat-critique runs Phases 1-5 plus mandatory post-synthesis steps (5.5, 5.6). Dispatcher ambiguity rules do not override direct invocation; raise scope concerns after synthesis.
Report-only by default. $goat-critique make X shorter = critique only; $goat-critique ... then apply it = critique first, apply after gate. See Constraints for mutation and apply rules.
goat-critique runs only full delegated mode: Phases 1-5, 5.5 meta-audit, 5.6 outcome capture, three critique sub-agents, one meta-agent. Lighter-mode suggestions are the failure this design prevents.
Intake checklist:
.goat-flow/learning-loop/footguns/ and .goat-flow/learning-loop/lessons/; record explicit misses instead of broad-loading buckets..goat-flow/logs/critiques/ has a same-artifact slug within 30 days, offer differential mode: A/B receive prior log + artifact diff; C stays cold. Phase 5 adds delta counts and [diff-of: <prior-uuid>].references/rubric-examples.md) and pass to each sub-agent's spawn directive.Spawn all three sub-agents in parallel using the host's real delegation mechanism.
Context varies intentionally - informational diversity catches more than tonal diversity.
Agents A and B both use the SKEPTIC/ANALYST/STRATEGIST combined lens. These three perspectives work as a unit - never split them into separate agents:
A/B must use all three perspectives. Agent C is the control group: same fields, no per-lens quota; use N/A - fresh-eyes scope only when a lens adds nothing.
Context split:
| Agent | Reads | Does NOT read |
|---|---|---|
| A (Risk) | artifact + architecture.md + targeted INDEX-first footgun/lesson hits + rubric | git history, config.yaml |
| B (Alternatives) | artifact + architecture.md + git log --oneline -20 + config.yaml + rubric | footguns, lessons |
| C (Fresh Eyes) | artifact + rubric ONLY | everything else (isolation enforced) |
Full directives: references/sub-agent-directives.md.
Each sub-agent normally returns 3-7 evidence-backed findings plus required lens fields, severity, evidence, confidence, Proof class, rubric dimensions, overall assessment, and one strength. Zero findings are valid only through the reference pack's clean-result attestation after one documented second pass.
Lens coverage: A/B must analyse every lens and C must probe assumptions/readability. A lens with no supported issue gets one documented second pass, then records No supported finding instead of inventing one. See the reference pack.
Execute in this order:
1. Context leak scan. Grep Agent C output for .goat-flow/, goat-*, architecture.md, config.yaml, or project-specific namespace references. Only flag references absent from Agent C's input. Untraceable match = CONTEXT LEAK; discard and re-spawn stricter. Framework-self exemption: for artifacts inside .goat-flow/, skills/goat-*, or a goat-flow instruction file, skip .goat-flow/ and goat-* term scans. Check only structural navigation leaks: file paths, config keys, or architecture sections absent from the input.
1b. Completeness gate. Verify each sub-agent returned required fields in either its findings or a clean-result attestation (Evidence reviewed: through Residual uncertainty: and strength per the reference pack, plus B's unconditional ranked alternative). Incomplete → re-spawn once.
2. Classify each finding: Consensus (≥2 agents, severity within ±1), Split (≥2 agents, severity differs ≥2 levels or explicit reject vs blocking), Unique (one agent only). Silence is not a dismiss; treat as Unique.
3. Score each sub-agent's critique on Grounding, Specificity, Actionability, Coverage, and Calibration.
4. Verify sub-agent dimension coverage. Skim each agent's findings, or a clean attestation's Rubric coverage: entries; confirm each claimed dimension has substantive content or named evidence. Demote unsubstantiated claims. Use orchestrator-verified dimensions as input to step 5.
5. Compute rubric coverage gates. A dimension is addressed by a substantive finding or step-4-verified attestation coverage. Unaddressed mandatory → auto-generate HIGH coverage-gap finding; optional → MEDIUM.
6. Spot-check OBSERVED claims. For each finding marked OBSERVED, re-read the cited file + semantic anchor or proof artifact. Findings that fail spot-check get tagged [evidence-gap: spot-check failed]; Phase 3 decides retract or upgrade.
7. Label control group deltas. For fresh-eyes-only findings, orchestrator assigns: CONTEXT DRIFT (wrong due to missing context), READABILITY GAP (valid for any reader), or CONTEXT-LIMITED (may be valid, cannot fully evaluate).
Early exit: If Phase 2 yields zero split findings and zero unique HIGH/CRITICAL findings, skip Phase 3. Note "no disputes - full consensus" in output and proceed to Phase 4.
If splits + unique HIGH/CRITICAL exceed the cross-examination budget (max 3 cross-exam agents total, 3 tool calls each - see Constraints), batch multiple disputes into a single agent prompt. Triage by severity - CRITICAL and HIGH first.
For each split finding, spawn a cross-exam agent: "Agent A says [X], Agent B says [Y]. Which is correct given the actual codebase?"
For unique HIGH/CRITICAL findings, spawn verification: "Only one critique raised [finding]. Genuine blind spot or false positive?"
Mark each: RESOLVED (with winner) / STILL DISPUTED / RETRACTED (false positive confirmed).
Persist before gate: Write Phase 1-3 results to .goat-flow/logs/critiques/<YYYY-MM-DD>-<HHMM>-<artifact-slug>-<rand5>.md - delegation evidence (ids/handles, calls/limit, unavailable markers), summaries, matrix, cross-exams. Runs even on Phase 3 early exit.
Before synthesising, present the unresolved items to the human conversationally.
Opener: Lead with a one-line summary of how many decisions are needed and their titles. Example: "3 decisions before synthesis: (1) SEC-01 severity, (2) remediation path, (3) attacker model scope."
Per-question format: Q[N]: [decision]? (A) [option] (B) [option] Default: [A/B]. Background: [1 sentence].
Compact table (3+ questions): | # | Decision | Option A (default) | Option B | Why |. Follow with: "Reply with numbers to override defaults; or approve to proceed."
Question types: (1) Disputes from Phase 3, (2) Trade-offs with two valid approaches, (3) Context drift findings - intentional vs oversight.
Closer: After all questions, end with: "Reply with your picks (e.g. 'A, B, go with defaults on the rest') or push back on any framing."
If questions exist: BLOCKING GATE - STOP and wait for human response. If no questions (full consensus, no trade-offs, no context drift): CHECKPOINT - note "no disputes - proceeding to synthesis" and continue.
Produce the prime critique. Lead with a Verdict block:
Resolved: N | Regressed: M | New: K | Unchanged: J vs prior critique)Then the full critique:
Open questions: Items with INFERRED-only evidence, inconclusive single-agent findings, or unvalidated assumptions go here - not as recommendations. Each open question states: confidence, evidence needed to resolve, revisit trigger.
Blind spot check: List unaddressed artifact sections, unmapped rubric aspects, and unread referenced files as "What Wasn't Critiqued." Must never be empty.
Proof Gate: Apply the Proof Gate (see Constraints) to every synthesised finding before inclusion. Every synthesised finding must carry proof class RUNTIME | CONTRACT-GREP | STATIC | NOT-REPRODUCED.
Phase 5.5 - Meta-audit. Spawn a lightweight meta-agent (budget: 2 tool calls, no context beyond the draft Phase 5 output). Audit the critique for internal consistency against the 10-point rubric in references/rubric-examples.md. If issues found, insert an ## Auto-Detected Issues block before presenting. Verdict block updated with Meta-score: N/100.
BLOCKING GATE: Present the synthesised critique (including Meta-score if 5.5 produced one). "Options: (A) apply, (B) dig deeper, (C) re-run, (D) close. Default: D." After plan critique, suggest /goat-plan.
Phase 5.6 - Outcome capture. After the human picks A/B/C/D, tag each surviving finding: accepted | rejected | deferred | partial. Defaults: A → accepted, D → deferred. Persist under ## Outcomes. Do not show ## Outcomes in the initial Phase 5 gate response.
Integration hooks. Populate from surviving findings when applicable:
for-goat-plan - milestone updates, reorderingfor-goat-debug - hypothesis seeds, evidence to capturefor-implementation - immediate fixes, deferred itemsEmpty sections collapsed to none.
The rubric determines what sub-agents evaluate. Match to artifact type. Dimensions marked [M] are mandatory (unaddressed → auto-HIGH coverage-gap finding); dimensions marked [O] are optional (unaddressed → auto-MEDIUM). Each rubric has a context map (A/B/C file assignments) in references/rubric-examples.md; Step 0 reads the selected map.
Plan: correctness against codebase [M], integration safety [M], sequencing quality [M], validation coverage [O], task specificity [O] Security assessment: threat model completeness [M], exploitability calibration [M], attack surface coverage [M], framework mitigation accuracy [O], data flow quality [O] Debug hypotheses: hypothesis diversity [M], evidence quality (OBSERVED vs INFERRED) [M], elimination rigour [M], confidence calibration [O], reproduction completeness [O] Review findings: severity calibration [M], diff coverage [M], pre-existing separation [M], false positive rate [O], cross-reference impact [O] Test strategy: coverage gaps [M], risk-proportionate depth [M], doer-verifier separation [O], manual test specificity [O], mock awareness [O] Architecture/refactor: blast radius accuracy [M], migration safety [M], backward compatibility [M], dependency impact [O], rollback feasibility [O] Generic (fallback): internal consistency [M], evidence grounding [M], scope completeness [M], feasibility [M], risk identification [M]. All dimensions mandatory for the fallback rubric. If using the generic rubric, state why no specific rubric matched and which was closest.
$goat-critique or /goat-critique invocation IS consent to spawn sub-agents and the full protocol. Do NOT ask again.references/sub-agent-directives.md. Incomplete → re-spawn once; if still incomplete, record sub-agent completeness limited.skill-preamble.md to every synthesised finding and preserve one proof class tag (RUNTIME | CONTRACT-GREP | STATIC | NOT-REPRODUCED) on each. Sub-agent reports are inputs to verify, not evidence to launder. Re-read applies to findings surviving to Phase 5 (typically 3-7 after Phase 3/4 filtering), not to all findings raised in Phase 1.Terse-first directive: Informational sections (Sub-Agent Comparison Matrix, Retracted Findings, What Wasn't Critiqued) default to terse: one sentence per bullet, no qualifiers, no closing offers. Gate prompts and evidence-tagged findings retain full detail.
Use this for the Phase 5 gate response. Omit ## Outcomes until Phase 5.6.
## Verdict <!-- includes Gate: BLOCK|CONCERNS|CLEAN + Meta-score -->
## Delegation Evidence <!-- ids/handles + tool-call counts or unavailable markers -->
## Critique Rubric
## Sub-Agent Comparison Matrix
## Sub-Agent Rankings
## Rubric Coverage Gaps
## Control Group Delta
## Validated Findings <!-- source pool for Recommended Changes; every finding includes proof class -->
## Cross-Examination Results
## Auto-Detected Issues <!-- from Phase 5.5 meta-audit, if any -->
## Retracted Findings
## Human Decisions
## Strengths
## Recommended Changes <!-- subset of Validated Findings; ordered by severity; each with concrete action and proof class -->
## Open Questions
## Integration Hooks <!-- for-goat-plan, for-goat-debug, for-implementation -->
## What Wasn't Critiqued
<!-- Phase 5.6 after human response: append ## Outcomes with per-finding accepted|rejected|deferred|partial -->