Skip to main content

benchmark-model-claims

Audits a claim about model or system performance against six integrity domains — test-set contamination, baseline pinning, seed and run variance, evaluator independence, metric selection, and whether the test set is large enough for the reported gap — and emits a 0–5 reliability score with named risk tags and a replication test. Use when a vendor post, paper or release note asserts a benchmark number: "91% on GSM8K", "outperforms GPT-4o", "SOTA on MMLU", "audit this benchmark claim". Not for grading a clinical or social-science trial (use `assess-study-bias`) and not for qualitative claims such as "better reasoning".

Jump to install

Source facts

Repository
radarist/structured-analytic-skills
Last source activity
August 18, 2026 at 13:10
Detected SKILL.md language
English
Stars
3
Forks
1

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.