Skip to main content

benchmark

Personal model benchmark. Replays the user's real recurring tasks (mined from their own Claude Code session history) against different models and effort levels in isolated headless sessions, grades the outputs blind against the user's own rubric, and produces a local HTML report plus a plain-verdict answer. Use whenever the user types /benchmark, asks to compare models ("opus vs fable", "opus 5 vs fable 5 on email triage, 3 trials"), asks if a new model is better or worth switching to, wants to compare effort levels (low vs high), or asks whether they can run a cheaper model/effort for the same result. Also handles "/benchmark mine" to build or refresh the test pack from history.

الانتقال إلى التثبيت

معلومات المصدر

المستودع
az9713/personal-model-benchmark
آخر نشاط في المصدر
٢٧ يوليو ٢٠٢٦ في ٠٢:٠٠
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
٠
التفرعات
٠

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.