Skip to main content

eval

Run one scripts-eval set — one `(target, question)` row from `experiments/scripts_eval/corpus.yaml` × 3 trials × one arm — including tester subagent dispatches + captures, plus (for arm C) judge subagent dispatches + records, then `summarize` + commit to the round's accumulator file. Use when the user says "run eval set", "eval", "scripts-eval", "round-NN set", or asks to execute a row of the corpus. Three arms: A (banned — rider forbids the antoine skills), B (directed — rider instructs use of antoine skills), C (organic — rider permits but doesn't direct). Two judge pairs: A-vs-B ("do the skills help when used") and A-vs-C ("do the skills get adopted organically"). `judge prepare --pair AB|AC` selects the pair.

インストールへ移動

ソース情報

リポジトリ
agentculture/antoine
ソースの最終更新活動
2026年5月26日 18:29
検出された SKILL.md の言語
英語
スター
2
フォーク
0

インストール方法

デフォルトでは、最初にソースを確認する Prompt が選択されています。直接コマンドに切り替えるか、ローカルコピーをダウンロードすることもできます。

ソースファイルを確認

インストールを決める前に、SKILL.md と SkillsMP に表示されている付属ファイルをお読みください。