| name | skill-eval-builder |
| description | Use when the user wants to measure or set up evals/checks for one of their skills โ how fast it is, whether its output is valid, whether it fires when expected, or whether its opening classification/routing gate labels inputs correctly. |
Skill Eval Builder
Overview
Set up a small, real eval for a skill: an evals/ folder next to it with a few
real cases, a runnable script, and a scorecard. It measures what scripts can't pin
down โ does the skill fire when it should (and stay quiet when it shouldn't),
is its output valid, is it within a time budget.
Core principle: an eval is a folder, not a framework. Keep it that small.
Inputs
- Target skill โ path or name (the dir with its
SKILL.md).
- A few real cases โ should-fire and shouldn't-fire prompts + real inputs; help the user find them if needed.
- Which dimensions matter โ default invocation + validation; add duration/others if they fit.
Steps
- Read the target skill โ its
SKILL.md, references, and scripts, so you know what it does and what it produces.
- Pick dimensions that fit it, from
references/eval-dimensions.md. Workflow skill โ invocation + duration; capability skill โ validation-heavy.
- Gather 3โ5 real cases โ should-fire and shouldn't-fire prompts plus real inputs. Always include at least one quiet case.
- Define "good" per case โ a
fired_when side-effect check, a validate command, and a budget_s from a real baseline (see references/eval-anatomy.md).
- Scaffold the artifact next to the skill:
python3 "${CLAUDE_SKILL_DIR}/scripts/scaffold-evals.py" <target-skill-dir>
Then fill in evals/cases.md with the cases from steps 3โ4.
- Check, then run the baseline:
python3 <target-skill-dir>/evals/run.py --dry-run
python3 <target-skill-dir>/evals/run.py
- Report the scorecard and where it lives.
Testing a classification gate
If the skill opens with a classifier (routes feature vs bug, or decides whether to continue), test that gate:
python3 "${CLAUDE_SKILL_DIR}/scripts/scaffold-evals.py" <target> --classify
- Fill
evals/classify-cases.md with the labels, the gate's own instruction, and real labeled examples (e.g. the last 10 tracker tickets + their existing labels as ground truth). See references/eval-anatomy.md.
python3 <target>/evals/classify.py โ an accuracy + confusion scorecard.
Output format
The scorecard (dimensions ร cases) plus the saved artifact path:
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโฌโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโ
โ Case โ Fires? โ Valid? โ Duration vs budget โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโผโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโค
โ "add tests for X" โ โ
โ โ
โ 40s / 60s โ
โ
โ "refactor Y" (shouldn't fire) โ โ
quiet โ โ โ โ โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโดโโโโโโโโโโโดโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโ
Saved to <target-skill>/evals/ (cases.md, run.py, results.md).
Guidelines
- Prefer an observable side-effect for firing (a file the skill produces) over grepping prose.
- Runs are not byte-identical โ a skill eval runs a model; that variance is what you measure. Don't fake determinism, and don't eval what a deterministic script already guarantees (unit-check the script instead).
- Keep it to 3โ5 real cases with at least one quiet case. See
references/eval-anatomy.md.
Files
scripts/scaffold-evals.py โ creates evals/{cases.md, run.py, results.md}, or the classify-* set with --classify.
scripts/run-evals.py โ runs cases via claude -p, times each, checks firing + validation, prints/saves the scorecard. --dry-run parses without spending tokens.
scripts/classify-evals.py โ classifies labeled examples via claude -p, reports accuracy + confusion. --dry-run parses without tokens.
references/eval-dimensions.md โ the dimension menu + when each applies.
references/eval-anatomy.md โ cases format, the side-effect pattern, anti-patterns.
If a script can't run here (needs python3 and the claude CLI, or a different OS): don't abandon the task โ an eval is just a cases.md plus a small runner, so run the cases and tally the scorecard with whatever tools are available.