Skip to main content

model-evaluation-specialist

Advanced model evaluation covering LLM benchmarks, evaluation frameworks (lm-evaluate-harness, HELM, RAGAS), leaderboard interpretation, custom metrics design, human evaluation protocols, automated LLM-as-judge patterns, and evaluation pipeline architecture for both traditional ML and generative AI systems. Use when the user asks about model evaluation specialist, model evaluation specialist best practices, or needs guidance on model evaluation specialist implementation. Do NOT use when the user needs a different specialized skill or is asking about an unrelated technology domain.

Aller à l'installation

Informations de source

Dépôt
FerroxLabs/wayland
Dernière activité de la source
7 juin 2026 à 16:09
Langue détectée de SKILL.md
anglais
Étoiles
569
Forks
107

Options d'installation

Le prompt qui vérifie d'abord la source est sélectionné par défaut. Vous pouvez passer à une commande directe ou télécharger une copie locale.

Vérifiez les fichiers source

Lisez SKILL.md et les fichiers associés affichés par SkillsMP avant de décider de l'installer.