Skip to main content
Run any Skill in Manus
with one click

orq-evaluator-alignment

Stars3
Forks0
UpdatedJuly 15, 2026 at 10:14

Align, calibrate, or improve an existing binary Pass/Fail LLM-as-a-judge (orq evaluator) so its verdicts match human judgment. Use when the user wants to "align my evaluator", "improve my eval", "annotate an evaluator", "find ambiguous cases", or "build an annotation queue" — i.e. they have a boolean judge (evaluator = LLM-as-a-judge) that disagrees with human labels or is inconsistent. Measures judge self-consistency (flip-rate) via repeated runs, surfaces the most ambiguous datapoints for human annotation, rewrites the judge prompt from the labels, and creates the new evaluator only after the human approves. If the evaluator ID isn't given, ask for it after triggering. Do NOT use to build an evaluator from scratch (use orq-build-evaluator), to fix failures with prompt tweaks (use orq-optimize-prompt), or for non-boolean judges.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

File Explorer
31 files
SKILL.md
readonly