eval-review
Review the latest enhance_notes eval results with scoring analysis, qualitative output review, and recommendations
التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.
القائمة
Review the latest enhance_notes eval results with scoring analysis, qualitative output review, and recommendations
التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.
استنادا إلى تصنيف SOC المهني
| name | eval-review |
| description | Review the latest enhance_notes eval results with scoring analysis, qualitative output review, and recommendations |
| user-invocable | true |
Review the latest eval results and produce a structured analysis.
.AI/EVAL_REVIEW_GUIDE.md for context — focus on: test case descriptions, known failure modes, research principles, and history of what's been tried..AI/results/summary.md for the score matrix, dimension averages, bullet stats, and ratio tables.Work through each of these, then produce a written review.
From summary.md:
Read results/<Test Case>.md for:
For each output you read, ask:
From the ratio tables:
Compare bullet growth vs word growth: if words grow faster than bullets, bullets are getting longer; if bullets grow faster, the model is proliferating.
Flag any case where the same variant's scores differ by >0.15 across runs. Read both outputs to understand what changed.
Any new variant must not regress on dense cases:
For each non-baseline variant, classify as:
Structure your review as:
NEVER include personal details in any git-tracked output (including EVAL_REVIEW_GUIDE.md). No real names, meeting content, transcript excerpts, or company-specific context. Keep references abstract — e.g. "a recurring standup" not the specific meeting name, "a management thread" not the people involved.
Run a full iteration of the enhance_notes prompt improvement loop — hypothesize, design, test, review
Offline CLI for replaying debug recordings through the audio pipeline for testing and parameter tuning. Use this skill whenever working with the replay tool, debug recordings, golden transcripts, WER scoring, parameter sweeps, or audio pipeline tuning. Also use when building, running, or debugging the replay binary.