원클릭으로
benchmark-critic
Evaluate benchmark design for fair baselines, measurable outcomes, leakage, cherry-picking, and reproducibility.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
메뉴
Evaluate benchmark design for fair baselines, measurable outcomes, leakage, cherry-picking, and reproducibility.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
SOC 직업 분류 기준
| name | benchmark-critic |
| description | Evaluate benchmark design for fair baselines, measurable outcomes, leakage, cherry-picking, and reproducibility. |
| license | MIT |
Review benchmark ideas or results for credibility.
Look for:
Recommend the smallest benchmark that can show useful signal without pretending to prove more than it does.
Audit an answer or implementation against primary sources before recommending changes.
Review a large question by staying inside one assigned slice, then support synthesis across several focused shard answers.
Check whether a change is ready for a public release candidate across packaging, docs, tests, and rollback risk.
Review code, docs, or plans for correctness risks, missing evidence, unclear tradeoffs, and untested assumptions.