Skip to main content

microsoft/ACESEvals

SkillsMP 已收集 microsoft/ACESEvals 中的 5 个 Skill。打开任一 Skill 可查看来源和详情。

最近记录的来源活动
SkillsMP 收录数据更新
已收集 skills
5
GitHub 星标
7
GitHub Forks
1

已展示 5 / 5 个已收集 Skill。

职业分类
未分类
描述

Guide for debugging inspect_ai evaluation failures, score issues, and model behavior. Use this when eval results are unexpected, scores are wrong, scoring fails, or model output appears corrupted.

原文语言:英语

更新
职业分类
未分类
描述

Guide for running SABER inspect_ai evaluations locally. Use this when asked to run, re-run, or configure an inspect eval for any SABER domain.

原文语言:英语

更新
职业分类
软件开发工程师
描述

Comprehensive guide for analyzing SABER evaluation results — model comparison, agent architecture comparison, domain-specific analysis, and cross-domain aggregate analysis. Use this when asked to analyze eval results, compare models, generate visualizations,…

原文语言:英语

更新
职业分类
网络与计算机系统管理员
描述

Guide for monitoring running SABER evaluations, checking progress, and managing eval batches. Use this when asked to monitor evals, check progress, produce a status report, or manage concurrent eval runs. Also covers Docker health and resource management.

原文语言:英语

更新
职业分类
软件质量保证分析师与测试员
描述

Guide for parsing and analyzing inspect_ai .eval log files. Use this when asked to interpret eval results, extract tool calls, find scores, or investigate agent behavior from .eval logs.

原文语言:英语

更新
已展示 5 / 5 个已收集 Skill。