一键导入
benchmark-critic
Evaluate benchmark design for fair baselines, measurable outcomes, leakage, cherry-picking, and reproducibility.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Evaluate benchmark design for fair baselines, measurable outcomes, leakage, cherry-picking, and reproducibility.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Audit an answer or implementation against primary sources before recommending changes.
Review a large question by staying inside one assigned slice, then support synthesis across several focused shard answers.
Check whether a change is ready for a public release candidate across packaging, docs, tests, and rollback risk.
Review code, docs, or plans for correctness risks, missing evidence, unclear tradeoffs, and untested assumptions.
| name | benchmark-critic |
| description | Evaluate benchmark design for fair baselines, measurable outcomes, leakage, cherry-picking, and reproducibility. |
| license | MIT |
Review benchmark ideas or results for credibility.
Look for:
Recommend the smallest benchmark that can show useful signal without pretending to prove more than it does.