Skip to main content

benchmark

Personal model benchmark. Replays the user's real recurring tasks (mined from their own Claude Code session history) against different models and effort levels in isolated headless sessions, grades the outputs blind against the user's own rubric, and produces a local HTML report plus a plain-verdict answer. Use whenever the user types /benchmark, asks to compare models ("opus vs fable", "opus 5 vs fable 5 on email triage, 3 trials"), asks if a new model is better or worth switching to, wants to compare effort levels (low vs high), or asks whether they can run a cheaper model/effort for the same result. Also handles "/benchmark mine" to build or refresh the test pack from history.

跳到安装

来源信息

仓库
az9713/personal-model-benchmark
最近来源活动
2026年7月27日 02:00
检测到的 SKILL.md 语言
英语
星标
0
分支
0

安装方式

默认使用会先检查来源的 Prompt;你也可以切换为直接命令,或下载本地副本。

检查来源文件

决定是否安装前,请先阅读 SKILL.md,以及 SkillsMP 当前展示的配套文件。