Skip to main content

benchmark

Personal model benchmark. Replays the user's real recurring tasks (mined from their own Claude Code session history) against different models and effort levels in isolated headless sessions, grades the outputs blind against the user's own rubric, and produces a local HTML report plus a plain-verdict answer. Use whenever the user types /benchmark, asks to compare models ("opus vs fable", "opus 5 vs fable 5 on email triage, 3 trials"), asks if a new model is better or worth switching to, wants to compare effort levels (low vs high), or asks whether they can run a cheaper model/effort for the same result. Also handles "/benchmark mine" to build or refresh the test pack from history.

Zur Installation springen

Quellinformationen

Repository
az9713/personal-model-benchmark
Letzte Quellaktivität
27. Juli 2026 um 02:00
Erkannte Sprache von SKILL.md
Englisch
Sterne
0
Forks
0

Installationsoptionen

Standardmäßig ist der Prompt ausgewählt, der zuerst die Quelle prüft. Sie können zu einem direkten Befehl wechseln oder eine lokale Kopie herunterladen.

Quelldateien prüfen

Lesen Sie SKILL.md und alle von SkillsMP angezeigten Begleitdateien, bevor Sie sich für eine Installation entscheiden.