Skip to main content

ml-benchmark-evaluation

Rigorous methodology for evaluating ML models on established benchmarks. Covers proper train/val/test splits, baseline verification from original papers, exact metric formula discrepancies, data-leak detection checklist, multi-seed robustness, and honest reporting templates. Use when claiming to beat published baselines, writing methods papers, or auditing existing results.

Jump to install

Source facts

Repository
synthetic-sciences/openscience
Last source activity
July 4, 2026 at 06:25
Detected SKILL.md language
English
Stars
3,362
Forks
454

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.