Skip to main content

braintrust-monitor-production-evals

Design or audit online evaluation of live LLM and agent traffic: trace sampling, online scoring coverage, alert thresholds and ownership, drift and failure-slice monitoring, incident review, and the pipeline that routes production failures back into the offline dataset. Use for questions about scoring production traces, monitoring quality after launch, catching regressions in the wild, alert fatigue, scorer drift, or closing the loop from incident to eval item. Do not use to design offline experiments or to interpret a controlled comparison.

Jump to install

Source facts

Repository
braintrustdata/eval-library
Last source activity
August 17, 2026 at 20:49
Detected SKILL.md language
English
Stars
12
Forks
2

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.