Skip to main content

benchmark-model-claims

Audits a claim about model or system performance against six integrity domains — test-set contamination, baseline pinning, seed and run variance, evaluator independence, metric selection, and whether the test set is large enough for the reported gap — and emits a 0–5 reliability score with named risk tags and a replication test. Use when a vendor post, paper or release note asserts a benchmark number: "91% on GSM8K", "outperforms GPT-4o", "SOTA on MMLU", "audit this benchmark claim". Not for grading a clinical or social-science trial (use `assess-study-bias`) and not for qualitative claims such as "better reasoning".

설치로 이동

소스 정보

저장소
radarist/structured-analytic-skills
최근 소스 활동
2026년 8월 18일 13:10
감지된 SKILL.md 언어
영어
스타
3
포크
1

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.