occupation
unclassified
description
Monitor and interpret evaluation runs of the eval-harness framework. Use when asked to watch, interpret, debug or diagnose an eval run, its logs, its scores or its results.
updated
Menu
SkillsMP has collected 3 skills from ScottRBK/eval-harness. Open a skill to review its source and details.
Showing 3 of 3 collected skills.
Monitor and interpret evaluation runs of the eval-harness framework. Use when asked to watch, interpret, debug or diagnose an eval run, its logs, its scores or its results.
Create eval-harness evaluations. Use when the user wants a new evaluation for a CLI coding agent (Claude Code, OpenCode, Copilot, Codex, Pi, Cursor, Grok).
Run evaluations with the eval-harness framework. Use when asked to run, execute, benchmark or compare CLI coding agents (Claude Code, OpenCode, Copilot, Codex, Pi, Cursor, Grok) on existing evals.