Skip to main content

eval

Run one scripts-eval set — one `(target, question)` row from `experiments/scripts_eval/corpus.yaml` × 3 trials × one arm — including tester subagent dispatches + captures, plus (for arm C) judge subagent dispatches + records, then `summarize` + commit to the round's accumulator file. Use when the user says "run eval set", "eval", "scripts-eval", "round-NN set", or asks to execute a row of the corpus. Three arms: A (banned — rider forbids the antoine skills), B (directed — rider instructs use of antoine skills), C (organic — rider permits but doesn't direct). Two judge pairs: A-vs-B ("do the skills help when used") and A-vs-C ("do the skills get adopted organically"). `judge prepare --pair AB|AC` selects the pair.

الانتقال إلى التثبيت

معلومات المصدر

المستودع
agentculture/antoine
آخر نشاط في المصدر
٢٦ مايو ٢٠٢٦ في ١٨:٢٩
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
٢
التفرعات
٠

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.