Audit an answer or implementation against primary sources before recommending changes.
Evaluate benchmark design for fair baselines, measurable outcomes, leakage, cherry-picking, and reproducibility.
Review a large question by staying inside one assigned slice, then support synthesis across several focused shard answers.
Check whether a change is ready for a public release candidate across packaging, docs, tests, and rollback risk.
Review code, docs, or plans for correctness risks, missing evidence, unclear tradeoffs, and untested assumptions.