Skip to main content

microsoft/ACESEvals

SkillsMP は microsoft/ACESEvals から 5 件の skill を収集しています。skill を開くとソースと詳細を確認できます。

記録された最新のソース活動
SkillsMP カタログ更新
収集済み skills
5
GitHub スター
7
GitHub フォーク
1

収集済み skill 5 件中 5 件を表示しています。

職業分類
未分類
説明

Guide for debugging inspect_ai evaluation failures, score issues, and model behavior. Use this when eval results are unexpected, scores are wrong, scoring fails, or model output appears corrupted.

原文の言語: 英語

更新
職業分類
未分類
説明

Guide for running SABER inspect_ai evaluations locally. Use this when asked to run, re-run, or configure an inspect eval for any SABER domain.

原文の言語: 英語

更新
職業分類
ソフトウェア開発者
説明

Comprehensive guide for analyzing SABER evaluation results — model comparison, agent architecture comparison, domain-specific analysis, and cross-domain aggregate analysis. Use this when asked to analyze eval results, compare models, generate visualizations,…

原文の言語: 英語

更新
職業分類
ネットワーク・コンピュータシステム管理者
説明

Guide for monitoring running SABER evaluations, checking progress, and managing eval batches. Use this when asked to monitor evals, check progress, produce a status report, or manage concurrent eval runs. Also covers Docker health and resource management.

原文の言語: 英語

更新
職業分類
ソフトウェア品質保証アナリスト・テスター
説明

Guide for parsing and analyzing inspect_ai .eval log files. Use this when asked to interpret eval results, extract tool calls, find scores, or investigate agent behavior from .eval logs.

原文の言語: 英語

更新
収集済み skill 5 件中 5 件を表示しています。