Claude Code style skill for benchmark research. Use when the user asks to analyze one paper's benchmarks, datasets, metrics, experiment tables/figures, baselines, or related work; or asks to survey a direction and find usable benchmarks for evaluation. Helper scripts fetch paper context/assets and Claude Code performs the semantic extraction.
2026-04-24