Skip to main content

agent-evaluation

Design reproducible evaluations for AI agents with representative task sets, explicit rubrics, appropriate graders, baselines, regression gates, and failure analysis. Use when defining agent quality, comparing prompts or models, validating a release, measuring tool-use reliability, investigating regressions, or deciding whether an agent is ready for production.

Jump to install

Source facts

Repository
seb1n/awesome-ai-agent-skills
Last source activity
August 9, 2026 at 18:55
Detected SKILL.md language
English
Stars
151
Forks
32

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.