Skip to main content

agent-eval

Reproducible evaluation harness for coding agents, prompts, and skills — head-to-head comparison with pass rate, cost, time, and consistency metrics captured in git worktrees. Use when comparing coding agents (Claude Code, Aider, Codex), regression-testing your own skills after changes, measuring variance across repeated runs of the same prompt, validating a model upgrade before adopting it, or A/B testing two versions of a prompt. Trigger on "compare Claude Code vs aider", "is my skill still passing", "run regression evals on this skill", "how much variance does this prompt have", "did the model upgrade regress anything". Use when this capability is needed.

跳到安装

来源信息

仓库
tomevault-io/skills-registry
最近来源活动
2026年5月23日 22:30
检测到的 SKILL.md 语言
英语
星标
0
分支
0

安装方式

默认使用会先检查来源的 Prompt;你也可以切换为直接命令,或下载本地副本。

检查来源文件

决定是否安装前,请先阅读 SKILL.md,以及 SkillsMP 当前展示的配套文件。