Skip to main content
Manusで任意のスキルを実行
ワンクリックで

loop-eval

スター1
フォーク0
更新日2026年7月22日 10:30

Design success criteria and a small eval set for an agentic loop — so the loop knows when it's done, and so YOU can measure whether the loop is any good. Use this whenever the user asks how to know if their agent/loop is working, how to define or sharpen success criteria, how to write evals/tests/graders for an agent, how to measure agent reliability, or wants pass/fail checks for a long-running task. Also use as the verification step when building a loop with loop-engineering. Covers verifiable criteria, code-based vs model-based graders, reliability metrics (pass@k vs pass^k), and balanced positive/negative cases. For a new unscoped loop, loop-spec owns success-definition intake first. Mining failures from specified run artifacts is led by loop-retro. Whole-harness audits are led by loop-review. This skill owns a standalone eval artifact once scope is settled. Routing tests for this suite are owned by loop-review.

インストール

Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。

ファイルエクスプローラー
3 ファイル
SKILL.md
readonly