Skip to main content

evals

Stars9
Forks5
UpdatedMay 2, 2026 at 15:24

Evaluation skill for AI output: apply measurable criteria to text/code/ research/policy. Supports four eval types: rubric-based (1-5 scale), pairwise (A vs B), checklist (binary), and red-team pass (uses thinking-redteam). Output: score + reasoning + improvement suggestions. [WHAT] Structured way to answer "is this good enough?". Uses domain rubrics: chronicle quality, memo rigor, academic precision, OSINT credibility. Eval is a feedback mechanism, not a gate. [WHEN] Use when: evaluate, score, judge, "is this good enough?", "rate this", quality check, compare A vs B. NOT for: fact-checking (use fact-check), academic peer review (use academic-opponent). [LANGUAGE] English and other languages; matches input.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

SKILL.md
readonly