Skip to main content

llm-evaluation

Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.

Jump to install

Source facts

Repository
wshobson/agents
Last source activity
July 7, 2026 at 16:17
Detected SKILL.md language
English
Stars
38,856
Forks
4,139

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.