Skip to main content

eval

Run one scripts-eval set — one `(target, question)` row from `experiments/scripts_eval/corpus.yaml` × 3 trials × one arm — including tester subagent dispatches + captures, plus (for arm C) judge subagent dispatches + records, then `summarize` + commit to the round's accumulator file. Use when the user says "run eval set", "eval", "scripts-eval", "round-NN set", or asks to execute a row of the corpus. Three arms: A (banned — rider forbids the antoine skills), B (directed — rider instructs use of antoine skills), C (organic — rider permits but doesn't direct). Two judge pairs: A-vs-B ("do the skills help when used") and A-vs-C ("do the skills get adopted organically"). `judge prepare --pair AB|AC` selects the pair.

跳到安装

来源信息

仓库
agentculture/antoine
最近来源活动
2026年5月26日 18:29
检测到的 SKILL.md 语言
英语
星标
2
分支
0

安装方式

默认使用会先检查来源的 Prompt;你也可以切换为直接命令,或下载本地副本。

检查来源文件

决定是否安装前,请先阅读 SKILL.md,以及 SkillsMP 当前展示的配套文件。