Skip to main content
تشغيل أي مهارة في Manus
بنقرة واحدة

benchmarks

النجوم٣
التفرعات٦
آخر تحديث٥ يوليو ٢٠٢٦ في ٢٠:٠٢

Review past benchmark runs and plan the next ones for the local model (Ouro / Σ₀ coder) — per-benchmark leaderboard tables placing the local stack against web-validated public SOTA on HumanEval, MBPP, SWE-bench, LongMemEval and honesty/HaluEval evals, plus a prioritized "run next" plan. Use whenever the user types `/benchmarks` or `!benchmarks`, or asks to "check the benchmarks", "how's the local model doing on evals", "plan the next benchmark runs", "where do we stand on HumanEval/SWE-bench/HaluEval", "compare Ouro to the council", "what should we benchmark next", or "benchmark the local model". Trigger even when the user names only one benchmark (e.g. "check our SWE-bench number") — reviewing one mark in the context of the whole ledger is this skill. Do NOT use it to grade the whole app (that's report-card), to review a PR diff (that's code-review), or to actually train/serve a model.

التثبيت

التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.

مستكشف الملفات
2 ملفات
SKILL.md
readonly