Guide for debugging inspect_ai evaluation failures, score issues, and model behavior. Use this when eval results are unexpected, scores are wrong, scoring fails, or model output appears corrupted.
microsoft/ACESEvals
SkillsMP has collected 5 skills from microsoft/ACESEvals. Open a skill to review its source and details.
- Latest recorded source activity
- SkillsMP catalog refreshed
- skills collected
- 5
- GitHub stars
- 7
- GitHub forks
- 1
Skills in this repository
Showing 5 of 5 collected skills.
Guide for running SABER inspect_ai evaluations locally. Use this when asked to run, re-run, or configure an inspect eval for any SABER domain.
Comprehensive guide for analyzing SABER evaluation results — model comparison, agent architecture comparison, domain-specific analysis, and cross-domain aggregate analysis. Use this when asked to analyze eval results, compare models, generate visualizations,…
Guide for monitoring running SABER evaluations, checking progress, and managing eval batches. Use this when asked to monitor evals, check progress, produce a status report, or manage concurrent eval runs. Also covers Docker health and resource management.
Guide for parsing and analyzing inspect_ai .eval log files. Use this when asked to interpret eval results, extract tool calls, find scores, or investigate agent behavior from .eval logs.