Skip to main content
Run any Skill in Manus
with one click

llm-benchmark

Stars18
Forks3
UpdatedMarch 23, 2026 at 14:16

Guide for designing, running, and interpreting LLM benchmark experiments — prompt ablation, statistical analysis of pass rates, per-turn interaction metrics, and data-leakage prevention. Use this skill when the user is: running benchmark tests against LLM prompts or configurations, comparing prompt variants (A/B testing prompts), analyzing benchmark results for statistical significance, designing test suites for LLM behavior, investigating per-turn LLM interaction quality, or asking whether sample sizes are sufficient. Also use when the user mentions ablation, pass rate, confidence intervals, Fisher exact test, or prompt optimization in a benchmarking context.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

File Explorer
2 files
SKILL.md
readonly