一键导入
evaluation-judge
Scores competing AI, colens, or prompt outputs against a declared rubric with evidence.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Scores competing AI, colens, or prompt outputs against a declared rubric with evidence.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Creator-side battle comparing the blog outline and YouTube script for the same topic.
Battle template comparing Claude and OpenAI on the same code review task using a four-axis rubric.
Compares single-lens and workflow-based finance report explanations. This is not financial advice.
Compares two models on creator planning using the YouTube content workflow.
Battle template comparing AI legal-review outputs on identical contract text. Analysis only — NOT legal advice.
Compares legal-adjacent document review outputs. This is not legal advice and must be reviewed by a qualified lawyer.
| name | evaluation-judge |
| description | Scores competing AI, colens, or prompt outputs against a declared rubric with evidence. |
Use this lenser when a battle or comparison needs a consistent rubric-based judgment.
The lenser may judge generated outputs. It must not hide uncertainty or choose a winner without evidence.
Return a score table, short rationale per criterion, winner, and residual uncertainty.