Skip to main content

eval-anova

Run Design-of-Experiments (DoE) evaluations with ANOVA over a matrix of agent configurations — comparing models, thinking-effort levels, prompts, or other factors across shared test cases, with repeated-measures / mixed-effects statistics that account for case difficulty plus a cost/quality Pareto view. Use whenever the user wants to compare models or configurations on an eval, decide which model or config is best, sweep or grid factors, run replications, or check whether a difference in eval scores is statistically significant (F, p, effect size) — even if they don't say "ANOVA" or "DoE". Also use when an eval.yaml has a matrix block or the user asks to fan an eval out across configurations.

Aller à l'installation

Informations de source

Dépôt
opendatahub-io/agent-eval-harness
Dernière activité de la source
13 août 2026 à 14:32
Langue détectée de SKILL.md
anglais
Étoiles
40
Forks
43

Options d'installation

Le prompt qui vérifie d'abord la source est sélectionné par défaut. Vous pouvez passer à une commande directe ou télécharger une copie locale.

Vérifiez les fichiers source

Lisez SKILL.md et les fichiers associés affichés par SkillsMP avant de décider de l'installer.