Skip to main content

eval-anova

Run Design-of-Experiments (DoE) evaluations with ANOVA over a matrix of agent configurations — comparing models, thinking-effort levels, prompts, or other factors across shared test cases, with repeated-measures / mixed-effects statistics that account for case difficulty plus a cost/quality Pareto view. Use whenever the user wants to compare models or configurations on an eval, decide which model or config is best, sweep or grid factors, run replications, or check whether a difference in eval scores is statistically significant (F, p, effect size) — even if they don't say "ANOVA" or "DoE". Also use when an eval.yaml has a matrix block or the user asks to fan an eval out across configurations.

الانتقال إلى التثبيت

معلومات المصدر

المستودع
opendatahub-io/agent-eval-harness
آخر نشاط في المصدر
١٣ أغسطس ٢٠٢٦ في ١٤:٣٢
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
٤٠
التفرعات
٤٣

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.