Skip to main content
在 Manus 中运行任何 Skill
一键导入

orq-evaluator-alignment

星标3
分支0
更新时间2026年7月15日 10:14

Align, calibrate, or improve an existing binary Pass/Fail LLM-as-a-judge (orq evaluator) so its verdicts match human judgment. Use when the user wants to "align my evaluator", "improve my eval", "annotate an evaluator", "find ambiguous cases", or "build an annotation queue" — i.e. they have a boolean judge (evaluator = LLM-as-a-judge) that disagrees with human labels or is inconsistent. Measures judge self-consistency (flip-rate) via repeated runs, surfaces the most ambiguous datapoints for human annotation, rewrites the judge prompt from the labels, and creates the new evaluator only after the human approves. If the evaluator ID isn't given, ask for it after triggering. Do NOT use to build an evaluator from scratch (use orq-build-evaluator), to fix failures with prompt tweaks (use orq-optimize-prompt), or for non-boolean judges.

安装

用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。

文件资源管理器
31 个文件
SKILL.md
readonly