Skip to main content
Run any Skill in Manus
with one click

eval-development

Stars0
Forks2
UpdatedJuly 16, 2026 at 20:56

Guide for adding new benchmarks to the agentm-eval framework and running experiments through the unified `agentm eval` CLI. Covers the adapter interface, experiment lifecycle, result schema, ClickHouse linking, and output conventions. Use whenever creating a new eval adapter, modifying an existing benchmark integration, writing code under contrib/evals/src/agentm_eval/, discussing experiment management, or when the user mentions "add a benchmark", "new eval", "ๆŽฅๅ…ฅ่ฏ„ไผฐ", "ๆ–ฐbenchmark", "ๅฎž้ชŒ็ฎก็†", "eval adapter". Also trigger when reviewing eval-related code or debugging experiment output structure.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

SKILL.md
readonly