원클릭으로
audit-loop-run
Use when asked to assess loop effectiveness, audit goal achievement, or detect phantom success.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
메뉴
Use when asked to assess loop effectiveness, audit goal achievement, or detect phantom success.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
SOC 직업 분류 기준
| name | audit-loop-run |
| description | Use when asked to assess loop effectiveness, audit goal achievement, or detect phantom success. |
Audit whether a configured loop's execution actually achieved its stated goal — checking artifact mutations, threshold contracts, structural defects (phantom convergence, degenerate gates, rubric drift, sub-loop verdict laundering), and producing ranked improvement proposals.
If loop_name argument is provided, resolve the most recent run folder:
ls -d .loops/.history/*-<loop_name>/ 2>/dev/null | sort | tail -1
<loop_name>." and stop.LATEST_RUN_ID (the compact timestamp prefix, e.g. 2026-03-19T204149).Otherwise, enumerate candidate loops:
ll-loop list --running --json
Filter to status one of "running", "interrupted", "failed", "timed_out", "awaiting_continuation". Sort by updated_at descending.
Note: ll-loop list --running --json output does not include instance_id — entries with the same loop_name are indistinguishable at this level.
loop_name values: Use AskUserQuestion to let the user pick:
Multiple loops found. Select one to assess:
[1] <loop_name_1> — <status> — last updated <updated_at>
[2] <loop_name_2> — <status> — last updated <updated_at>
...
loop_name (multiple instances): follow up with ll-loop status <loop_name> --json to retrieve per-instance detail (instance_id, pid, log_file, events_file), then use AskUserQuestion to present instance-level disambiguation:
Multiple instances of '<loop_name>' found. Select one to assess:
[1] <instance_id_1> — <status> — PID <pid> — last updated <updated_at>
[2] <instance_id_2> — <status> — PID <pid> — last updated <updated_at>
...
Before loading or analyzing anything, confirm the run artifacts exist and are non-empty. This applies to every path that reaches this step — auto-resolved runs, directly-supplied run IDs/folders, and running-loop selections alike.
RUN_DIR=".loops/.history/<LATEST_RUN_ID>-<loop_name>"
if [ ! -s "$RUN_DIR/events.jsonl" ] || [ ! -f "$RUN_DIR/state.json" ]; then
echo "MISSING_RUN"
fi
MISSING_RUN (or RUN_DIR does not exist): report
Run '<LATEST_RUN_ID>-<loop_name>' not found or empty — refusing to audit.
and stop. Do not emit a verdict, state-transition trace, captured
outputs, improvement proposals, or any other section. An audit of a run whose
events.jsonl/state.json cannot be read is a fabrication, not an audit —
the only honest output is the refusal above.events.jsonl/state.json. If a tool call returns empty or errors, treat
that as absence of evidence, not an invitation to confabulate.Only once the gate passes, proceed.
Load the fully-materialized FSM:
ll-loop show <loop_name> --resolved --json
This returns FSMLoop.to_dict() JSON with always-present keys name, initial, states, and conditionally description, context (threshold keys live here), max_steps, parameters, commands.
Load the event history. If the user supplied --tail N, use that directly. Otherwise, auto-scale to load all events (--tail 0):
# Derive total event count from the archive (line count = event count)
TOTAL_EVENTS=$(wc -l .loops/.history/<LATEST_RUN_ID>-<loop_name>/events.jsonl | awk '{print $1}')
# Use user-supplied tail if provided, else 0 (all events)
EFFECTIVE_TAIL=<tail_arg_or_0>
ll-loop history <loop_name> [<LATEST_RUN_ID>] --json --tail ${EFFECTIVE_TAIL}
If EFFECTIVE_TAIL is greater than 0 and less than TOTAL_EVENTS, emit a truncation notice before proceeding:
ℹ️ Loaded last <EFFECTIVE_TAIL> of <TOTAL_EVENTS> events — fault analysis covers a partial window.
If either command fails, report the error and stop.
From the FSM context flat dict, scan for threshold keys:
target_pass_rate, pass_threshold, quality_threshold, readiness_thresholdoutcome_threshold, reward_target, target_score, min_per_category, adversarial_capAlso scan each state's action text and evaluate.prompt text for ${context.<key>} interpolation patterns to detect threshold references embedded in prompts.
Build the success contract: list of {key, value, source} entries where source is "context", "action", or "evaluate.prompt".
If no contract entries are found, note: "No threshold contract detected — loop uses implicit success criteria."
Identify artifact paths the loop touches. Look in:
context.prompt_file, context.system_file, context.output_file, context.run_dir and similar path-like context keysaction text for file path patterns (prompts/, data/, .issues/, image.svg, manifest.json, examples.json)For each identified artifact path, check mutation evidence:
# Check if file was modified in recent git history
git log --oneline -5 -- <artifact_path>
# Check current diff
git diff HEAD -- <artifact_path>
For issue-based loops, inspect frontmatter:
ll-issues show <id> --json
Also check in-memory captures in .loops/.history/<run_id>-<loop_name>/state.json under captured dict (schema: {capture_variable_name: {output, stderr, exit_code, duration_ms}} — keys are capture variable names from capture: declarations, not state names). For step-level capture output in events.jsonl, read action_complete.output_preview.
Quote every .output value verbatim when citing it; do not infer "sentinel" or "placeholder" labels — the interpolation engine emits no numeric markers (only \x00ESCAPED\x00, an internal placeholder that is never present in captured output).
Re-use the history loaded in Step 2 to identify fault signals using the fault-signal subset of /ll:debug-loop-run Step 3 (the BUG-class anomalies that broke the run). Note: /ll:debug-loop-run Step 3 also classifies effectiveness signals (iter-1 convergence without apply, degenerate gate, stub action) — those are out of scope for audit-loop-run Phase 1, since this step only synthesizes fault evidence into the scorecard. Include the verbatim fault signal list in the scorecard output.
Key signals to flag (fault subset only):
exit_code != 0, non-intentional)evaluate.verdict == "error" on the last evaluate before loop_complete) — single-occurrence terminating evaluator error (eval_error_termination); distinct from "Evaluate failures" which covers verdict == "fail" 3+ timesverdict == "fail", 3+ occurrences on the same state)throttle_stop = loop halted; throttle_hard = loop redirected via on_throttle_hard).output value matches ^\d{2,7}\b (a bare PID prefix) and the action text for that state contains $$( or $$[A-Za-z_] (same pattern as _OVERESCAPED_SHELL_RE in validation.py:121), flag as over-escaped-shell-pid-corruption (MR-9) and recommend removing the extra $, never adding more escaping.Using the history loaded in Step 2 and the artifact evidence from Step 4, run the shallow-iteration heuristic:
1. Count tool calls: Count the number of action_complete events in the loaded history. Call this TOOL_CALL_COUNT.
2. Identify auxiliary mutations: Before trusting git diff HEAD, check whether the primary artifact path (or the run's working directory, e.g. context.run_dir) is gitignored:
git check-ignore <primary_path>
Not ignored (exit code 1, or no primary path is under version control at all): use the git diff HEAD evidence collected in Step 4 as before — list all files that were created or modified and are not in the primary artifact path set (the paths identified in Step 4: context.prompt_file, context.output_file, context.run_dir, and similar path-like context keys). Call this count AUX_MUTATION_COUNT.
Ignored (exit code 0): git diff HEAD is structurally blind to this path, so a git-derived AUX_MUTATION_COUNT of 0 cannot be trusted. Fall back to a filesystem mutation scan scoped to the run's working directory, anchored on the run-start timestamp (events[0].ts from the history loaded in Step 2):
# GNU find (Linux) — accepts an ISO timestamp string directly
find <run_dir> -type f -newermt "<run_start_ts>"
# BSD find (macOS default) — -newermt is unsupported; use a touched marker file instead
touch -d "<run_start_ts>" /tmp/run_start_marker && find <run_dir> -type f -newer /tmp/run_start_marker
Count the resulting file list as AUX_MUTATION_COUNT.
Neither signal available (e.g. the run directory has already been cleaned up): do not default to 0. Report AUX_MUTATION_COUNT as unknown and route the heuristic result to unknown in step 4 below (skip the warning rather than asserting a false positive).
3. Check for diff_stall corroboration: Scan evaluate events in the history. For each, check whether the resolved FSM (loaded in Step 2) has evaluate.type == "diff_stall" for that state and the recorded verdict is "stall" or "no". Call this DIFF_STALL_PRESENT (true/false).
4. Apply threshold (default threshold: 30):
IF AUX_MUTATION_COUNT == "unknown":
result = "unknown" # no git or filesystem evidence available — skip, don't guess
ELIF TOOL_CALL_COUNT > 30 AND AUX_MUTATION_COUNT == 0:
IF DIFF_STALL_PRESENT:
result = "corroborated" # both heuristic and diff_stall agree
ELSE:
result = "warning" # heuristic alone
ELSE:
result = "clear"
The default threshold of 30 action_complete events is intentionally conservative — most well-structured loops either produce auxiliary artifacts or converge within this budget. Loops that burn more than 30 iterations without creating helper structure are iterating without building.
5. Emit finding when result is "warning" or "corroborated":
⚠ Shallow-iteration: <TOOL_CALL_COUNT> action_complete events with no auxiliary file mutations
outside the primary artifact path (<primary_paths>).
[Corroborated by diff_stall evaluator verdict in state '<state_name>'.]
Remediation: add intermediate artifact-write states; break monolithic iteration into
smaller sub-tasks that each produce a named helper file.
Pass result and TOOL_CALL_COUNT to the scorecard in Step 6.
Before accepting budget-exhaustion as a root cause, compute the budget utilization ratio:
STEPS_CONSUMED=$(jq '[.[] | select(.event == "loop_complete")] | last | .iterations // 0' \
.loops/.history/<LATEST_RUN_ID>-<loop_name>/events.jsonl)
MAX_STEPS=$(ll-loop show <loop_name> --resolved --json | jq '.max_steps // .max_iterations // 100')
If STEPS_CONSUMED / MAX_STEPS < 0.3, reject budget-exhaustion as the primary root cause — the loop consumed less than 30% of its budget, so it did not run out of steps.
Note: there is no steps_consumed field in state.json; derive STEPS_CONSUMED from loop_complete.iterations in events.jsonl.
Before determining the verdict, check whether the run wrote a summary.json to its run directory:
SUMMARY_FILE=".loops/.history/<LATEST_RUN_ID>-<loop_name>/summary.json"
If the file exists, extract the claimed-outcome counters (closed, implemented, failed, decomposed). The success token varies by loop — auto-refine-and-implement / sprint-refine-and-implement emit closed (verified terminal closure, ENH-2385), while rn-implement / general-task emit implemented. Use whichever success counter the loop reports as the claimed-success signal:
closed > 0 / implemented > 0 (or any equivalent success token) is present0 (or key absent) — the run honestly reports it produced nothingENH-2404 — parked-issue visibility (auto-refine-and-implement / autodev): if present, also read skipped_breakdown (an object keyed by reason, e.g. {"decomposed": 1, "refine_failed": 0, "low_readiness": 4}), gate_blocked (issues parked by the learning-gate, ENH-2402 — previously invisible here), and parked_rate ((skipped + not_closed + gate_blocked) / input_size). parked_rate is a visibility signal, not a pass/fail gate — interpret it via skipped_breakdown: a high rate dominated by decomposed is healthy (the run is legitimately fanning out into children), while one dominated by refine_failed / low_readiness is a genuine quality signal worth flagging in the report. These three keys are additive; older summary.json files (pre-ENH-2404) will lack them — treat their absence as "no breakdown data available" rather than an error, and fall back to the plain skipped count.
ENH-2533 — per-issue + learning-followup visibility (rn-implement): if present, also read per_issue (an array of {id, outcome, reason?, pre_scores?, post_scores?, convergence?} records aggregated from the run's subloop_outcome_<ID>.txt sidecars — one per issue the queue touched) and learning_followups (an array of {id, targets, remedy} records aggregated from learning_unproven_<ID>.txt sidecars, where remedy is /ll:explore-api <targets>). Cite specific parked-issue IDs from per_issue in the verdict rationale instead of bucketed counters — e.g. "ENH-400 parked with MANUAL_REVIEW_RECOMMENDED, BUG-401 parked with LEARNING_GATE_BLOCKED" — so the audit reproduces the operator's screen-readable rationale. These are additive; older summary.json files (pre-ENH-2533) will lack both keys — fall back to the learning_gate_blocked scalar counter and the per-record subloop_outcome_<ID>.txt sidecars directly. Malformed per-issue sidecars surface in summary_warnings.txt (not summary.json); check there only when per_issue is shorter than the run's tally of parked IDs.
ENH-2601 — post-implementation verify verdict (auto-refine-and-implement / sprint-refine-and-implement): if present, also read verify_verdict ("passed" / "failed" / "skipped" / "not_run") — the result of running project.test_cmd/lint_cmd once, after delegate and before finalize. This is advisory only: it does not gate the run's own verdict (a closed > 0 run can still report verify_verdict: "failed" if a regression slipped through). "skipped" means test_cmd was unconfigured; "not_run" means the verify state never executed (e.g. the resolved issue set was empty, or delegate crashed before verify could run). Flag verify_verdict: "failed" prominently in the report even when the closure verdict itself reads success — that combination is exactly the gap this field exists to surface. This key is additive; older summary.json files (pre-ENH-2601) will lack it — treat its absence as "no verify data available," not an error.
Determine the verdict using the terminal state from loop_complete event (terminated_by), the artifact/contract evidence from Step 4, and the claimed-success signal from Step 6a:
| Verdict | Condition |
|---|---|
met | Terminal reached AND all threshold contracts verified AND all expected artifact mutations occurred |
phantom | Terminal reached AND claimed success > 0 (or summary.json absent — loop provides no failure evidence) AND (artifacts unchanged OR threshold unverified — only model self-reported via llm_structured evaluator) |
honest-failure | Terminal reached AND summary.json present AND claimed success == 0 (implemented: 0, failed: N) AND no artifact mutation observed. The loop told the truth about its failure; the root cause is upstream (e.g. environment error, auth failure, misconfiguration). |
partial | Terminal reached AND some but not all contracts satisfied |
partial | terminated_by == "max_steps" AND max_steps_summary event present in JSONL (summary state ran; artifact written) |
degraded | Loop completed but metric trended downward vs baseline captured in state.json |
Output the structured scorecard block:
### Goal-vs-Outcome Scorecard
**Goal**: "<loop description or (no description provided)>"
**Contract**: <threshold keys and values, or "none detected">
**Artifacts checked**: <list of paths and mutation status>
**Phase 1 signals**: <fault signal count from Step 5, or "none">
**Shallow-iteration check**: `<warning | corroborated | clear | unknown>` (<TOOL_CALL_COUNT> tool calls, <AUX_MUTATION_COUNT> auxiliary mutations)
**Verdict**: `<met | phantom | honest-failure | partial | degraded>`
**Rationale**: <one paragraph explaining the verdict>
Skip this step if --no-rubric-audit flag is set.
For each state with evaluate.type: llm_structured, send a judge call comparing:
description textprompt textJudge prompt (single call per evaluator):
"Does this evaluator prompt operationalize the loop's stated goal? Loop goal: ''. Evaluator prompt: '<evaluate.prompt>'. Answer YES if the evaluator directly measures progress toward the stated goal, NO if it measures something unrelated or misaligned."
Flag as rubric drift if the judge answers NO. Include the evaluator's state name and a brief explanation.
Pattern reference: outer-loop-eval.yaml:generate_report state uses evaluate.type: llm_structured with min_confidence: 0.7.
For each state where loop: is set (sub-loop invocation), read on_yes and on_no from the FSM JSON output:
state.on_yes # child reached a terminal state
state.on_no # child did not reach terminal
Laundering defect: state.on_yes == state.on_no (after any ${context.*} interpolation). This means the parent loop treats child success and child failure identically — the child verdict is silently discarded.
ENH-2005 sidecar exemption: Before flagging, check whether the artifact-channel sidecar pattern is present. A state is exempt when all of the following hold:
action contains subloop_outcome_ — the child writes its real verdict to this artifact and the parent recovers it downstream.state.on_error is set and routes to a distinct state (not the shared classifier target) — ensuring an infrastructure crash is attributed separately, not collapsed into the generic failure path.When both conditions hold, do not flag as a laundering defect. Instead, note [mitigated — ENH-2005 artifact-channel sidecar: verdict recovered via subloop_outcome_ artifact, on_error routes to distinct crash state]. When on_error is also collapsed into the shared target, or the shared target does not read subloop_outcome_, flag as before — those cases are genuinely unsafe.
Flag each unmitigated laundering defect with:
loop: value)on_yes and on_no point to)Emit ranked proposals from the scorecard, rubric audit, and fault signals. Order: contract-level > rubric-level > state-level > structural.
For each proposal, include a concrete YAML diff where possible:
### Improvement Proposals
1. [contract] Add artifact mutation verification for `prompts/test.md`
Rationale: loop reached terminal without evidence of file mutation — possible phantom success
YAML diff:
states:
optimize:
+ capture: optimized_prompt
+ capture_file: "${context.prompt_file}"
2. [rubric] Align evaluator prompt with loop goal in state `refine_answers`
Rationale: evaluate.prompt checks Python syntax; description says "improve answer quality"
3. [state] Add `on_error` routing to state `check_quality`
Rationale: shell evaluator with no on_error silently routes failed runs to on_no
Before presenting proposals, check for existing issues:
grep -rl "<loop_name>" .issues/bugs/ .issues/enhancements/ .issues/features/ .issues/epics/ 2>/dev/null
Mark matches as DUPLICATE. Present only NEW proposals.
Skip this step if --skip-issue-creation or --auto flag is set (or if LL_NON_INTERACTIVE/DANGEROUSLY_SKIP_PERMISSIONS env vars are set, or --dangerously-skip-permissions is active). Print: ℹ️ Issue creation skipped (--skip-issue-creation / --auto) and stop.
Use AskUserQuestion to ask:
Create issues for these <N> proposals? [Y/n/select]
Y — create all
n — cancel
select — choose which to create (comma-separated numbers)
For each approved proposal, allocate an ID (ll-issues next-id) and write the issue file to the appropriate category dir. Stage each written file by its explicit path (git add "<issue-file-path>") — do not git add .issues/, which sweeps in unrelated untracked/modified files (BUG-1976).
Assessment complete for loop: <loop_name>
Verdict: `<met | phantom | honest-failure | partial | degraded>`
Rubric audit: <N evaluators checked, M flagged — or "skipped (--no-rubric-audit)">
Laundering check: <N sub-loop states checked, M flagged — or "no sub-loop states">
Shallow-iteration check: `<warning | corroborated | clear | unknown>` (<N> tool calls, <M> auxiliary mutations — or "below threshold")
Issues created: <N>
# Assess most recent interrupted loop
/ll:audit-loop-run
# Assess a specific loop
/ll:audit-loop-run apo-textgrad
# Limit history to 100 events
/ll:audit-loop-run apo-textgrad --tail 100
# Skip LLM rubric audit (cost gate)
/ll:audit-loop-run apo-textgrad --no-rubric-audit
When this skill emits an audit finding, verdict, or scorecard, cite evidence verbatim rather than re-summarizing — quoting is cheaper than paraphrasing and keeps the audit auditable:
IMPORTANT: For each condition you evaluate:
Do not assert a verdict without evidence. "The task appears complete" is not evidence.
Rewrite an issue's Implementation Steps, Acceptance Criteria, and Files to Modify in place from its own accumulated research findings, without appending or bulldozing human prose
Use when asked to audit documentation accuracy, coverage, or find documentation gaps.
Use when asked about project health, velocity, bug trends, or whether we're making progress.
Use when asked for an adversarial go/no-go review or whether an issue is worth implementing.
Use when asked to manually compact a session's memory, trigger session summarization, or reduce a long session's context footprint.
Use when asked to manually compact a session's memory, trigger session summarization, or reduce a long session's context footprint.