ワンクリックで
benchmark-critic
Evaluate benchmark design for fair baselines, measurable outcomes, leakage, cherry-picking, and reproducibility.
Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。
メニュー
Evaluate benchmark design for fair baselines, measurable outcomes, leakage, cherry-picking, and reproducibility.
Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。
Audit an answer or implementation against primary sources before recommending changes.
Review a large question by staying inside one assigned slice, then support synthesis across several focused shard answers.
Check whether a change is ready for a public release candidate across packaging, docs, tests, and rollback risk.
Review code, docs, or plans for correctness risks, missing evidence, unclear tradeoffs, and untested assumptions.
SOC 職業分類に基づく
| name | benchmark-critic |
| description | Evaluate benchmark design for fair baselines, measurable outcomes, leakage, cherry-picking, and reproducibility. |
| license | MIT |
Review benchmark ideas or results for credibility.
Look for:
Recommend the smallest benchmark that can show useful signal without pretending to prove more than it does.