Skip to main content

benchmark

Personal model benchmark. Replays the user's real recurring tasks (mined from their own Claude Code session history) against different models and effort levels in isolated headless sessions, grades the outputs blind against the user's own rubric, and produces a local HTML report plus a plain-verdict answer. Use whenever the user types /benchmark, asks to compare models ("opus vs fable", "opus 5 vs fable 5 on email triage, 3 trials"), asks if a new model is better or worth switching to, wants to compare effort levels (low vs high), or asks whether they can run a cheaper model/effort for the same result. Also handles "/benchmark mine" to build or refresh the test pack from history.

インストールへ移動

ソース情報

リポジトリ
az9713/personal-model-benchmark
ソースの最終更新活動
2026年7月27日 02:00
検出された SKILL.md の言語
英語
スター
0
フォーク
0

インストール方法

デフォルトでは、最初にソースを確認する Prompt が選択されています。直接コマンドに切り替えるか、ローカルコピーをダウンロードすることもできます。

ソースファイルを確認

インストールを決める前に、SKILL.md と SkillsMP に表示されている付属ファイルをお読みください。