Skip to main content
Ejecuta cualquier Skill en Manus
con un clic

llm-benchmark

Estrellas18
Forks3
Actualizado23 de marzo de 2026 a las 14:16

Guide for designing, running, and interpreting LLM benchmark experiments — prompt ablation, statistical analysis of pass rates, per-turn interaction metrics, and data-leakage prevention. Use this skill when the user is: running benchmark tests against LLM prompts or configurations, comparing prompt variants (A/B testing prompts), analyzing benchmark results for statistical significance, designing test suites for LLM behavior, investigating per-turn LLM interaction quality, or asking whether sample sizes are sufficient. Also use when the user mentions ablation, pass rate, confidence intervals, Fisher exact test, or prompt optimization in a benchmarking context.

Instalación

Instalar con Codex o Claude Copia este prompt, pégalo en Codex, Claude u otro asistente, y deja que revise la página de la skill y la instale por ti.

Explorador de archivos
2 archivos
SKILL.md
readonly