Skip to main content

reliability-benchmark

Run a first-party reliability benchmark comparing two or more LLMs on hard, trap-laden tasks through one identical harness, graded by reading and scored pass^k. Use when the user wants to benchmark, compare, or A/B models for reliability (not raw capability), test a new model release against a baseline, or reproduce the Kimi K3 vs Opus study. Triggers on "benchmark these models", "compare model X vs Y", "run the reliability benchmark", "test <model> against <model>", "is <model> actually as good as its scores".

Zur Installation springen

Quellinformationen

Repository
coleam00/kimi-k3-reliability-benchmark
Letzte Quellaktivität
21. Juli 2026 um 13:16
Erkannte Sprache von SKILL.md
Englisch
Sterne
1
Forks
2

Installationsoptionen

Standardmäßig ist der Prompt ausgewählt, der zuerst die Quelle prüft. Sie können zu einem direkten Befehl wechseln oder eine lokale Kopie herunterladen.

Quelldateien prüfen

Lesen Sie SKILL.md und alle von SkillsMP angezeigten Begleitdateien, bevor Sie sich für eine Installation entscheiden.