Skip to main content

validate-evaluator

Calibrate an LLM judge against human labels using data splits, TPR/TNR, and bias correction. Use after writing a judge prompt (write-judge-prompt) when you need to verify alignment before trusting its outputs. Do NOT use for code-based evaluators (those are deterministic; test with standard unit tests).

Jump to install

Source facts

Repository
hamelsmu/evals-skills
Last source activity
June 10, 2026 at 05:28
Detected SKILL.md language
English
Stars
1,664
Forks
170

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.