Guide for debugging inspect_ai evaluation failures, score issues, and model behavior. Use this when eval results are unexpected, scores are wrong, scoring fails, or model output appears corrupted.
원문 언어: 영어
메뉴
SkillsMP는 microsoft/ACESEvals에서 5개의 skill을 수집했습니다. skill을 열어 소스와 세부 정보를 확인하세요.
수집된 skill 5개 중 5개를 표시합니다.
Guide for debugging inspect_ai evaluation failures, score issues, and model behavior. Use this when eval results are unexpected, scores are wrong, scoring fails, or model output appears corrupted.
원문 언어: 영어
Guide for running SABER inspect_ai evaluations locally. Use this when asked to run, re-run, or configure an inspect eval for any SABER domain.
원문 언어: 영어
Comprehensive guide for analyzing SABER evaluation results — model comparison, agent architecture comparison, domain-specific analysis, and cross-domain aggregate analysis. Use this when asked to analyze eval results, compare models, generate visualizations,…
원문 언어: 영어
Guide for monitoring running SABER evaluations, checking progress, and managing eval batches. Use this when asked to monitor evals, check progress, produce a status report, or manage concurrent eval runs. Also covers Docker health and resource management.
원문 언어: 영어
Guide for parsing and analyzing inspect_ai .eval log files. Use this when asked to interpret eval results, extract tool calls, find scores, or investigate agent behavior from .eval logs.
원문 언어: 영어