Skip to main content
تشغيل أي مهارة في Manus
بنقرة واحدة

evaluate-submission

النجوم٦
التفرعات١
آخر تحديث٣ يونيو ٢٠٢٦ في ٢٠:٢٧

Run the Simulation Bench automated harness against a submission folder under /submissions and produce an evaluation_report.json plus a concise summary. Use this skill whenever the user asks to evaluate, score, grade, run the harness on, or check a submission to the Simulation Bench — phrasings like "evaluate submission X", "run the harness on Y", "score the latest run", "did model Z pass?", "check submissions/<folder>" should all trigger it. Also use when the user mentions a submission folder name that follows the date__benchmark__harness__model taxonomy, even if they don't say the word "evaluate".

التثبيت

التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.

SKILL.md
readonly