Skip to main content

benchmark-design

Designs reproducible ML/LLM evaluation benchmarks including metric selection (BLEU, ROUGE, pass@k), evaluation protocols (few-shot, zero-shot, chain-of-thought), dataset splits with contamination prevention, and statistical significance testing. Use when building or auditing evaluation suites, leaderboards, or model comparison frameworks. Do not use for training pipelines or data curation.

Aller à l'installation

Informations de source

Dépôt
merceralex397-collab/meta-skill-engineering
Dernière activité de la source
20 avril 2026 à 01:25
Langue détectée de SKILL.md
anglais
Étoiles
2
Forks
0

Options d'installation

Le prompt qui vérifie d'abord la source est sélectionné par défaut. Vous pouvez passer à une commande directe ou télécharger une copie locale.

Vérifiez les fichiers source

Lisez SKILL.md et les fichiers associés affichés par SkillsMP avant de décider de l'installer.