Skip to main content
Execute qualquer Skill no Manus
com um clique

llm-benchmark

Estrelas18
Forks3
Atualizado23 de março de 2026 às 14:16

Guide for designing, running, and interpreting LLM benchmark experiments — prompt ablation, statistical analysis of pass rates, per-turn interaction metrics, and data-leakage prevention. Use this skill when the user is: running benchmark tests against LLM prompts or configurations, comparing prompt variants (A/B testing prompts), analyzing benchmark results for statistical significance, designing test suites for LLM behavior, investigating per-turn LLM interaction quality, or asking whether sample sizes are sufficient. Also use when the user mentions ablation, pass rate, confidence intervals, Fisher exact test, or prompt optimization in a benchmarking context.

Instalação

Instalar com Codex ou Claude Copie este prompt, cole no Codex, Claude ou outro assistente e deixe que ele revise a página da skill e instale para você.

Explorador de arquivos
2 arquivos
SKILL.md
readonly