- name
- arm-evals
- description
- Arm a forge eval suite into a frozen build contract. Run when .forge/state.json shows phase=evals, after /plugin-forge:build-evals; re-run after any suite edit or triage grader-error to re-version and re-freeze. Four gates in order via forge-eval: reference (every task's reference solution passes its own graders), red (no-plugin baseline; tasks that pass bare carry no signal), metaeval (judge vs human agreement >= 0.85 against evals/labels/labels.jsonl), then sha256 FREEZE into .forge/freeze.json with suite_version vN and phase -> armed. Refuses to freeze uncalibrated judges, unresolved no-signal tasks, or a suite with pending change requests.
- allowed-tools
- Bash(forge-eval *), Bash(find evals *)
!`cat .forge/state.json 2>/dev/null || echo NO_STATE`
# arm-evals — four gates, then the freeze
If the first line of this skill rendered as literal backtick syntax instead of JSON or `NO_STATE`
(org policy `disableSkillShellExecution` blocks dynamic injection), read `.forge/state.json` with the
Read tool before doing anything else.
A drafted suite is not yet a contract. It becomes one only after proving three things — its graders
can recognize a correct solution, bare Claude cannot pass it, and its judges agree with a human —
and then being made tamper-evident by a cryptographic freeze. Run the four gates strictly in order;
each later gate is more expensive than the one before it, and an earlier failure invalidates later
results.
## State guard
Route on the injected state above. Refuse out-of-order execution and give the exact remediation
command — never "help anyway".
| Injected state | Action |
|---|---|
| `NO_STATE` | Refuse: no forge project in this directory. Remediation: run `/plugin-forge:forge` |
| `phase: interview` | Refuse: the interview is unfinished. Remediation: run `/plugin-forge:forge` |
| `phase: contract` | Refuse: no approved contract yet. Remediation: run `/plugin-forge:write-contract` |
| `phase: evals` | Proceed — this is the arming phase. If `evals/` does not exist yet, refuse instead: run `/plugin-forge:build-evals` |
| `phase: armed` | Proceed only as a RE-ARM (see "Re-arming"). If the tree hash still equals `.forge/freeze.json`'s sha256, report "already armed at <suite_version>, nothing to do" and stop |
| `phase: building` | Refuse: a build is in flight and the frozen suite must not change under it. Remediation: wait for `forge-build` to exit (or clear a crashed run with `forge-eval doctor`), then run `/plugin-forge:arm-evals` |
| `phase: smoke` or `phase: verify` | Proceed only as a RE-ARM driven by a triage grader-error finding or an approved change request. Warn first: re-arming regresses phase to `armed` and voids the current green — `/plugin-forge:build-loop` must run again |
| `phase: ship` | Proceed as a RE-ARM starting the next iteration cycle (new suite version); phase returns to `armed` |
Read `plugin_dir` from `.forge/state.json` — every gate command below needs it.
## Before gate 1
- If a previous run crashed or state looks inconsistent, run `forge-eval doctor` first.
- Check `.forge/eval-change-requests.md`: review every pending entry with the user, apply or reject
each, and record the decision in the file. Refuse to freeze over unreviewed requests.
- Cost honesty before starting: gate 2 executes the full suite headlessly (tasks x trials). Estimate
from the task count, tell the user, and pass `--max-cost-usd` on every run. Use `--dry-run` once to
validate wiring before spending anything, and `--jobs N` (max 8) to parallelize.
## Gate 1 — REFERENCE GREEN
Every task's `reference/` solution must pass that task's own graders, headlessly. For each suite in
`evals/registry.yaml`:
```
forge-eval reference --suite <suite> --plugin-dir <plugin_dir>
```
Doctrine: a task whose own reference solution scores 0% is a **broken task, not a hard task**. The
grader cannot recognize a known-correct answer, so it can never recognize the builder's. On any
failure, fix the task or the grader (with the user for judgment calls), then re-run this gate for
the affected tasks. Do not proceed to gate 2 with a red reference.
Run this gate first because it is the cheapest and because a broken grader corrupts every later
gate's evidence.
## Gate 2 — RED BASELINE
Run the full suite against the no-plugin target (`targets/no-plugin.yaml`, i.e. bare Claude with the
plugin absent):
```
forge-eval red --suite <suite> --plugin-dir <plugin_dir> --max-cost-usd <X>
```
- Every task SHOULD fail here. A task that PASSES without the plugin is flagged **"no signal"** — it
measures nothing the plugin adds. For each no-signal task, choose with the user:
- **Harden** it: strengthen the end state, add a second verification signal, tighten a tolerance,
or require an artifact only the plugin's workflow produces — while keeping the
two-experts-same-verdict bar. Every hardened task must re-pass gate 1.
- **Drop** it: delete the task if hardening would make it artificial.
- Store the baseline: the red run's `runs/<ts>/summary.json` is the with/without-plugin denominator
that verify later reports the improvement delta against. Never delete that run directory; cite its
path in the arming report.
- Trigger-suite tasks trivially show no activation under the no-plugin target; that is expected and
is not a no-signal finding — trigger precision/recall is measured later by verify under bare and
crowded targets.
## Gate 3 — JUDGE CALIBRATION
For every judge spec under `evals/judges/` referenced by any task grader:
1. **Labels first.** Calibration needs 10–20 human-labeled transcripts per judge dimension, drawn
from the gate 1 and gate 2 runs (`runs/<ts>/transcripts/`). If `evals/labels/labels.jsonl` is
missing or too thin, sample the transcripts, present them to the user for hand-labeling in the
scaffolded labels.jsonl format, and WAIT. **Never fill labels.jsonl by model judgment** — a
model-labeled golden set makes the agreement metric circular and voids this gate entirely.
2. **Measure agreement:**
```
forge-eval metaeval --plugin-dir <plugin_dir>
```
This reports per-judge agreement between the judge's verdicts and the human labels.
3. **Agreement >= 0.85** → the judge passes. Below 0.85, run the Align-Evals iteration loop:
- sort disagreements worst-first and read those transcripts alongside the judge's own
rationale/evidence output;
- add explicit **negative criteria** to the judge prompt for each observed false-pass pattern;
- insert the corrected cases into the judge prompt as **few-shot examples**;
- **re-version the judge spec** (bump its version marker so scores stay attributable to a judge
version), re-run metaeval, and **re-baseline** agreement per judge version;
- repeat until agreement reaches 0.85 or the judge is replaced.
4. Judges run on a pinned cheap model that is not the target model. Changing the pinned judge model
invalidates calibration — re-run metaeval after any model change.
**Refusal is mandatory here.** If any judge is uncalibrated (no labels, too few labels, or agreement
below 0.85), refuse to freeze and say exactly why: an uncalibrated judge silently converts
"measured" into "vibes" — its scores cannot bind a build loop, and a green result would be
unfalsifiable. Do not accept deleting the judge grader as a workaround: the PDR capability row that
judge verifies becomes unobservable, and unverifiable means unbuildable. The legitimate paths are
labeling the transcripts (preferred) or replacing the judge with a deterministic two-signal check —
which is a contract change requiring a change-request entry and a spec-critic re-review.
## Gate 4 — FREEZE
Only when gates 1–3 are green **against the exact current tree**. If any file under `evals/` changed
while running the gates (hardened tasks, iterated judges, added labels), re-run gate 1 for edited
tasks and gate 3 for edited judges before freezing.
1. **Determine the suite version N.** No `.forge/freeze.json` → `v1`. Otherwise recompute the tree
hash (step 3's pipeline): identical to the recorded sha256 means nothing changed (report "already
armed", stop); different means N = previous N + 1. Any suite edit auto-bumps N — that is what
keeps scores comparable across suite revisions.
2. **Bump the registry BEFORE hashing.** Update every alias in `evals/registry.yaml` to
`<suite>.<split>.v<N>`. The registry lives inside `evals/`, so editing it after computing the
hash would invalidate the freeze immediately.
3. **Compute the tree sha256** from the project root:
```
find evals -type f -print0 | LC_ALL=C sort -z | xargs -0 shasum -a 256 | shasum -a 256 | awk '{print $1}'
```
This pipeline is the canonical TREE HASH CONTRACT documented in
`${CLAUDE_PLUGIN_ROOT}/hooks/scripts/sweep-evals.sh` and mirrored by `tree_sha256` in the
harness (`scripts/harness/run.py`). Never substitute a different hashing recipe — the Stop-hook
sweep compares against exactly this value during the build, and a divergent recipe would either
block every stop or detect nothing.
4. **Write `.forge/freeze.json`** (it lives outside `evals/`, so it does not perturb the hash):
```json
{"sha256": "<tree-hash>", "suite_version": "v<N>", "frozen_at": "<ISO8601 UTC>"}
```
5. **Update `.forge/state.json`**: set `phase` to `"armed"` and `updated` to the current ISO8601
timestamp; preserve every other field.
6. **Holdout check:** confirm `evals-holdout/` exists OUTSIDE `evals/` (project root). It is
deliberately excluded from the freeze hash and is never seeded into the build worktree — it is
the anti-overfit tranche the builder must never see. If build-evals created no holdout, flag it
to the user before proceeding.
## Re-arming after any suite edit
Triggers: a triage report classifying a failure as **grader-error**; an approved entry in
`.forge/eval-change-requests.md`; any manual edit under `evals/` (labels included — they are part of
the hashed tree); a judge-model change.
Procedure: review pending change requests with the user, apply the edits, re-run the affected gates
(gate 1 for every edited task, gate 2 when tasks were hardened/added/dropped, gate 3 for any touched
judge), then gate 4 — the version bump happens automatically because the tree hash differs.
Re-arming from `smoke`, `verify`, or `ship` regresses phase to `armed`: prior green results were
measured against the old suite version and no longer bind.
## Refusal conditions (all hard)
Refuse to freeze while any of these holds, stating which gate failed, why the freeze would be
unsound, and the exact next action:
1. Any reference solution failing its own graders (gate 1).
2. Any unresolved no-signal task (gate 2).
3. Any judge with agreement below 0.85, or with missing/insufficient human labels (gate 3).
4. `labels.jsonl` populated by the model instead of the human — circular calibration (gate 3).
5. Unreviewed entries in `.forge/eval-change-requests.md`.
6. `phase: building` — never re-freeze under a live build.
## Handoff
Report the arming result on one line plus the next command, for example:
```
ARMED: suite_version=v2 sha256=3fa9c1d2e0b4... tasks=24 judges=2 (agreement 0.91, 0.88) red_baseline=runs/2026-08-02T14-11-05/
Next: /plugin-forge:build-loop
```
The suite is now the contract: during the build it is mechanically read-only (PreToolUse guard +
Stop-hook sweep, holding even under `--dangerously-skip-permissions`), and the only legitimate way
to change it is a change request plus a human re-running this skill.
Ver en GitHub