职业分类
未分类
描述
Monitor and interpret evaluation runs of the eval-harness framework. Use when asked to watch, interpret, debug or diagnose an eval run, its logs, its scores or its results.
原文语言:英语
更新
菜单
SkillsMP 已收集 ScottRBK/eval-harness 中的 3 个 Skill。打开任一 Skill 可查看来源和详情。
已展示 3 / 3 个已收集 Skill。
Monitor and interpret evaluation runs of the eval-harness framework. Use when asked to watch, interpret, debug or diagnose an eval run, its logs, its scores or its results.
原文语言:英语
Create eval-harness evaluations. Use when the user wants a new evaluation for a CLI coding agent (Claude Code, OpenCode, Copilot, Codex, Pi, Cursor, Grok).
原文语言:英语
Run evaluations with the eval-harness framework. Use when asked to run, execute, benchmark or compare CLI coding agents (Claude Code, OpenCode, Copilot, Codex, Pi, Cursor, Grok) on existing evals.
原文语言:英语