Skip to main content
Ejecuta cualquier Skill en Manus
con un clic

benchmarks

Estrellas3
Forks6
Actualizado5 de julio de 2026 a las 20:02

Review past benchmark runs and plan the next ones for the local model (Ouro / Σ₀ coder) — per-benchmark leaderboard tables placing the local stack against web-validated public SOTA on HumanEval, MBPP, SWE-bench, LongMemEval and honesty/HaluEval evals, plus a prioritized "run next" plan. Use whenever the user types `/benchmarks` or `!benchmarks`, or asks to "check the benchmarks", "how's the local model doing on evals", "plan the next benchmark runs", "where do we stand on HumanEval/SWE-bench/HaluEval", "compare Ouro to the council", "what should we benchmark next", or "benchmark the local model". Trigger even when the user names only one benchmark (e.g. "check our SWE-bench number") — reviewing one mark in the context of the whole ledger is this skill. Do NOT use it to grade the whole app (that's report-card), to review a PR diff (that's code-review), or to actually train/serve a model.

Instalación

Instalar con Codex o Claude Copia este prompt, pégalo en Codex, Claude u otro asistente, y deja que revise la página de la skill y la instale por ti.

Explorador de archivos
2 archivos
SKILL.md
readonly