Skip to main content

agent-eval

Reproducible evaluation harness for coding agents, prompts, and skills — head-to-head comparison with pass rate, cost, time, and consistency metrics captured in git worktrees. Use when comparing coding agents (Claude Code, Aider, Codex), regression-testing your own skills after changes, measuring variance across repeated runs of the same prompt, validating a model upgrade before adopting it, or A/B testing two versions of a prompt. Trigger on "compare Claude Code vs aider", "is my skill still passing", "run regression evals on this skill", "how much variance does this prompt have", "did the model upgrade regress anything". Use when this capability is needed.

설치로 이동

소스 정보

저장소
tomevault-io/skills-registry
최근 소스 활동
2026년 5월 23일 22:30
감지된 SKILL.md 언어
영어
스타
0
포크
0

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.