원클릭으로
evaluation-judge
Scores competing AI, colens, or prompt outputs against a declared rubric with evidence.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
메뉴
Scores competing AI, colens, or prompt outputs against a declared rubric with evidence.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
SOC 직업 분류 기준
| name | evaluation-judge |
| description | Scores competing AI, colens, or prompt outputs against a declared rubric with evidence. |
Use this lenser when a battle or comparison needs a consistent rubric-based judgment.
The lenser may judge generated outputs. It must not hide uncertainty or choose a winner without evidence.
Return a score table, short rationale per criterion, winner, and residual uncertainty.
Creator-side battle comparing the blog outline and YouTube script for the same topic.
Battle template comparing Claude and OpenAI on the same code review task using a four-axis rubric.
Compares single-lens and workflow-based finance report explanations. This is not financial advice.
Compares two models on creator planning using the YouTube content workflow.
Battle template comparing AI legal-review outputs on identical contract text. Analysis only — NOT legal advice.
Compares legal-adjacent document review outputs. This is not legal advice and must be reviewed by a qualified lawyer.