| name | eval-skills |
| description | Eval and improve a skill against golden cases — run the target skill blind in a fresh, context-free subagent on each example input, grade the artifact against the expected outcome, and let the gaps drive the edits. Use when the user wants to test/eval/improve/harden a skill, says "this skill keeps producing X / keeps missing Y", or hands a skill plus example input→expected-output pairs. Pairs with [write-skills](../write-skills/SKILL.md) (the authoring principles every fix obeys). |
Eval Skills
Treat a skill like a function under test. Feed it example inputs in a clean
room, check the artifacts against what good looks like, and let the failures
drive the edits. The eval is only honest if the run is blind: the agent
executing the skill must carry none of this conversation's context and must
never see the expected output. Leak either and you are teaching to the test.
Inputs you need — refuse without them
Confirm all three before spawning anything. If any is missing or
unresolvable, stop and tell the user exactly which one and what a good
version looks like. Do not invent cases, guess intent, or eval against a
fuzzy wish.
- Target skill — must resolve to a real
SKILL.md. If you can't find it,
list the skills you can see and ask which one they mean.
- At least one golden case — a concrete input the skill will actually
receive: a screenshot, a prompt, a file, a scene. "Improve write-spec"
with no input attached is not a case.
- The bar per case — the outcome a good artifact achieves and the smells
that would make it bad, not an exhaustive parts list. The skill's
judgment is what's under test, so do not pre-enumerate every
requirement — that turns the eval into a conformance check and stops testing
whether the skill decides well. "Sliced so each piece is independently
buildable and verifiable, at the granularity a competent practitioner would
pick — a lazy mega-slice and pointless over-splitting are both failures" is
a bar a judge can hold the work to; "slices it well" is too thin to grade
and a fixed list of expected slices is too prescriptive. State the bar and
the smells; let the judge apply them. The exception is a
skill that genuinely wants an exact task hit exactly — then the explicit
criteria the bar; match the bar's shape to the skill's nature, and if
you can't tell which it is, ask. If the user gives only a fuzzy wish with no
bar, draw the bar out of them and echo it back before spending agents.