Skip to main content

benchmark-model-claims

Audits a claim about model or system performance against six integrity domains — test-set contamination, baseline pinning, seed and run variance, evaluator independence, metric selection, and whether the test set is large enough for the reported gap — and emits a 0–5 reliability score with named risk tags and a replication test. Use when a vendor post, paper or release note asserts a benchmark number: "91% on GSM8K", "outperforms GPT-4o", "SOTA on MMLU", "audit this benchmark claim". Not for grading a clinical or social-science trial (use `assess-study-bias`) and not for qualitative claims such as "better reasoning".

Aller à l'installation

Informations de source

Dépôt
radarist/structured-analytic-skills
Dernière activité de la source
18 août 2026 à 13:10
Langue détectée de SKILL.md
anglais
Étoiles
3
Forks
1

Options d'installation

Le prompt qui vérifie d'abord la source est sélectionné par défaut. Vous pouvez passer à une commande directe ou télécharger une copie locale.

Vérifiez les fichiers source

Lisez SKILL.md et les fichiers associés affichés par SkillsMP avant de décider de l'installer.