Skip to main content
Jeden Skill in Manus ausführen
mit einem Klick

benchmarks

Sterne3
Forks6
Aktualisiert5. Juli 2026 um 20:02

Review past benchmark runs and plan the next ones for the local model (Ouro / Σ₀ coder) — per-benchmark leaderboard tables placing the local stack against web-validated public SOTA on HumanEval, MBPP, SWE-bench, LongMemEval and honesty/HaluEval evals, plus a prioritized "run next" plan. Use whenever the user types `/benchmarks` or `!benchmarks`, or asks to "check the benchmarks", "how's the local model doing on evals", "plan the next benchmark runs", "where do we stand on HumanEval/SWE-bench/HaluEval", "compare Ouro to the council", "what should we benchmark next", or "benchmark the local model". Trigger even when the user names only one benchmark (e.g. "check our SWE-bench number") — reviewing one mark in the context of the whole ledger is this skill. Do NOT use it to grade the whole app (that's report-card), to review a PR diff (that's code-review), or to actually train/serve a model.

Installation

Mit Codex oder Claude installieren Kopieren Sie diesen Prompt, fügen Sie ihn in Codex, Claude oder einen anderen Assistant ein und lassen Sie die Skill-Seite prüfen und installieren.

Datei-Explorer
2 Dateien
SKILL.md
readonly