Skip to main content

braintrust-validate-eval-scorer

Validate automated eval scorers and LLM judges against expert-reviewed reference data. Use to compare scorer output with human labels, calculate agreement (kappa, alpha) with uncertainty, inspect confusion by class and severity, analyze subgroup failures, test shortcut and gaming cases, propagate scorer error into headline numbers, document blind spots, and decide whether a scorer is fit for exploration, trend monitoring, or release gating. Do not use to create the initial scorer or to design the human review workflow.

Jump to install

Source facts

Repository
braintrustdata/eval-library
Last source activity
August 17, 2026 at 20:49
Detected SKILL.md language
English
Stars
12
Forks
2

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.