Skip to main content

microsoft/ACESEvals

SkillsMP has collected 5 skills from microsoft/ACESEvals. Open a skill to review its source and details.

Latest recorded source activity
SkillsMP catalog refreshed
skills collected
5
GitHub stars
7
GitHub forks
1

Showing 5 of 5 collected skills.

occupation
unclassified
description

Guide for debugging inspect_ai evaluation failures, score issues, and model behavior. Use this when eval results are unexpected, scores are wrong, scoring fails, or model output appears corrupted.

updated
occupation
unclassified
description

Guide for running SABER inspect_ai evaluations locally. Use this when asked to run, re-run, or configure an inspect eval for any SABER domain.

updated
occupation
Software Developers
description

Comprehensive guide for analyzing SABER evaluation results — model comparison, agent architecture comparison, domain-specific analysis, and cross-domain aggregate analysis. Use this when asked to analyze eval results, compare models, generate visualizations,…

updated
occupation
Network & Computer Systems Administrators
description

Guide for monitoring running SABER evaluations, checking progress, and managing eval batches. Use this when asked to monitor evals, check progress, produce a status report, or manage concurrent eval runs. Also covers Docker health and resource management.

updated
occupation
Software Quality Assurance Analysts & Testers
description

Guide for parsing and analyzing inspect_ai .eval log files. Use this when asked to interpret eval results, extract tool calls, find scores, or investigate agent behavior from .eval logs.

updated
Showing 5 of 5 collected skills.