Skip to main content

evals

星标9
分支5
更新时间2026年5月2日 15:24

Evaluation skill for AI output: apply measurable criteria to text/code/ research/policy. Supports four eval types: rubric-based (1-5 scale), pairwise (A vs B), checklist (binary), and red-team pass (uses thinking-redteam). Output: score + reasoning + improvement suggestions. [WHAT] Structured way to answer "is this good enough?". Uses domain rubrics: chronicle quality, memo rigor, academic precision, OSINT credibility. Eval is a feedback mechanism, not a gate. [WHEN] Use when: evaluate, score, judge, "is this good enough?", "rate this", quality check, compare A vs B. NOT for: fact-checking (use fact-check), academic peer review (use academic-opponent). [LANGUAGE] English and other languages; matches input.

安装

用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。

SKILL.md
readonly