| name | braintrust-design-human-eval-review |
| description | Design or audit human evaluation workflows and golden datasets for LLM applications and agents. Use to set up expert review, select review cases, write reviewer instructions, assign raters, capture rationales and confidence, measure inter-rater agreement with kappa or alpha, adjudicate disagreements, and preserve reviewed examples with provenance as a versioned reference set. Do not use to elicit the criteria or rubric in the first place, to validate a scorer once reference labels exist, or to implement the scorer. |
Design human review and the golden set
Contract: references/interaction-contract.md. Calibration, templates, provenance: references/review-workflow.md.
Trigger
- "Set up human review." / "Build a golden dataset." / "How should we adjudicate?"
- A rubric that exists and now needs applying at scale by people.
- A judge that needs an anchor before it can be trusted.
Do
- Identify the expertise required and the decision the labels will support. Draft the review
form before asking about reviewer count.
- If no rubric or criteria exist yet, stop and elicit them from experts first. This skill
applies criteria; it does not invent them.
- Select cases deliberately: representative items, plus the ambiguous region where systems
differ, plus every severe failure class.
- Collect a short rationale and confidence with every label — the rationales are where the
rubric gets sharp.
- Have raters judge independently before conferring, then adjudicate on the record. A panel
deferring to whoever speaks first discards the benefit of a panel.
- Report agreement and adjudicate before the set calibrates anything, then preserve it as a
versioned dataset with per-item provenance and a refresh trigger.
Avoid
- Do not treat majority opinion as ground truth without the relevant expertise.