Skip to main content

eval

Run one scripts-eval set — one `(target, question)` row from `experiments/scripts_eval/corpus.yaml` × 3 trials × one arm — including tester subagent dispatches + captures, plus (for arm C) judge subagent dispatches + records, then `summarize` + commit to the round's accumulator file. Use when the user says "run eval set", "eval", "scripts-eval", "round-NN set", or asks to execute a row of the corpus. Three arms: A (banned — rider forbids the antoine skills), B (directed — rider instructs use of antoine skills), C (organic — rider permits but doesn't direct). Two judge pairs: A-vs-B ("do the skills help when used") and A-vs-C ("do the skills get adopted organically"). `judge prepare --pair AB|AC` selects the pair.

Jump to install

Source facts

Repository
agentculture/antoine
Last source activity
May 26, 2026 at 18:29
Detected SKILL.md language
English
Stars
2
Forks
0

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.