Skip to main content

evals

스타9
포크5
업데이트2026년 5월 2일 15:24

Evaluation skill for AI output: apply measurable criteria to text/code/ research/policy. Supports four eval types: rubric-based (1-5 scale), pairwise (A vs B), checklist (binary), and red-team pass (uses thinking-redteam). Output: score + reasoning + improvement suggestions. [WHAT] Structured way to answer "is this good enough?". Uses domain rubrics: chronicle quality, memo rigor, academic precision, OSINT credibility. Eval is a feedback mechanism, not a gate. [WHEN] Use when: evaluate, score, judge, "is this good enough?", "rate this", quality check, compare A vs B. NOT for: fact-checking (use fact-check), academic peer review (use academic-opponent). [LANGUAGE] English and other languages; matches input.

설치

Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.

SKILL.md
readonly