Skip to main content

evals

Objective eval metrics via code/model/human graders with pass@k/pass^k scoring. USE WHEN eval, evaluate, test agent, benchmark, verify behavior, regression test, capability test, run eval, compare models, compare prompts, create judge, create use case, view results, failure to task, suite manager, transcript capture, trial runner.

Jump to install

Source facts

Repository
danielmiessler/Personal_AI_Infrastructure
Last source activity
March 15, 2026 at 22:10
Detected SKILL.md language
English
Stars
15,967
Forks
2,205

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.