ワンクリックで
my-eval-plan
Design evaluation plans for AI/LLM features: datasets, scorers, baselines, success criteria, and regression strategy.
Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。
メニュー
Design evaluation plans for AI/LLM features: datasets, scorers, baselines, success criteria, and regression strategy.
Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。
SOC 職業分類に基づく
Address pending PR review feedback through verified triage, small fix phases, implementation, validation, and evidence-backed replies. Manual invocation only.
Autonomous iteration loop for a measurable goal: review, ideate, modify, verify, keep or rollback, repeat until interrupted or capped.
Monitor a PR's CircleCI pipeline, diagnose failures, apply scoped fixes, push when requested, and continue until green or blocked.
Stage and commit changes in logical groups using the project's git message style, without mixing unrelated user changes.
Create or update a GitHub PR with concise description, review guidance, triggered specialty reviews, and focus areas.
Build a daily work brief from yesterday/off-hours activity plus today's Linear, Calendar, Gmail, and Notion context; update Notion and produce standup/checklist.
| name | my-eval-plan |
| description | Design evaluation plans for AI/LLM features: datasets, scorers, baselines, success criteria, and regression strategy. |
Design a practical evaluation strategy before or alongside AI feature work.
Read ~/.claude/rules/question-policy.md when available. Use ~/.agents/rules/ under Codex. For full planning procedure, read references/protocol-index.md.
Return scorer definitions, dataset plan, baseline targets, instrumentation needs, and implementation checklist.