Skip to main content
تشغيل أي مهارة في Manus
بنقرة واحدة

loop-eval

النجوم١
التفرعات٠
آخر تحديث٢٢ يوليو ٢٠٢٦ في ١٠:٣٠

Design success criteria and a small eval set for an agentic loop — so the loop knows when it's done, and so YOU can measure whether the loop is any good. Use this whenever the user asks how to know if their agent/loop is working, how to define or sharpen success criteria, how to write evals/tests/graders for an agent, how to measure agent reliability, or wants pass/fail checks for a long-running task. Also use as the verification step when building a loop with loop-engineering. Covers verifiable criteria, code-based vs model-based graders, reliability metrics (pass@k vs pass^k), and balanced positive/negative cases. For a new unscoped loop, loop-spec owns success-definition intake first. Mining failures from specified run artifacts is led by loop-retro. Whole-harness audits are led by loop-review. This skill owns a standalone eval artifact once scope is settled. Routing tests for this suite are owned by loop-review.

التثبيت

التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.

مستكشف الملفات
3 ملفات
SKILL.md
readonly