Skip to main content

evals

النجوم٩
التفرعات٥
آخر تحديث٢ مايو ٢٠٢٦ في ١٥:٢٤

Evaluation skill for AI output: apply measurable criteria to text/code/ research/policy. Supports four eval types: rubric-based (1-5 scale), pairwise (A vs B), checklist (binary), and red-team pass (uses thinking-redteam). Output: score + reasoning + improvement suggestions. [WHAT] Structured way to answer "is this good enough?". Uses domain rubrics: chronicle quality, memo rigor, academic precision, OSINT credibility. Eval is a feedback mechanism, not a gate. [WHEN] Use when: evaluate, score, judge, "is this good enough?", "rate this", quality check, compare A vs B. NOT for: fact-checking (use fact-check), academic peer review (use academic-opponent). [LANGUAGE] English and other languages; matches input.

التثبيت

التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.

SKILL.md
readonly