Skip to main content
Run any Skill in Manus
with one click

pre-registering-eval-study

Stars2
Forks0
UpdatedJune 17, 2026 at 16:52

Locks the hypothesis, falsification criteria, stopping rules, and power justification of an empirical evaluation BEFORE any data is collected or observed. Use when designing an LLM eval, an A/B experiment, a clinical or observational study, an ML model comparison, or any decision that will rest on a test statistic. Produces a one-page pre-registration document committed to a versioned location prior to the first run, so that HARKing (hypothesizing after results are known), p-hacking, optional-stopping, and outcome-switching cannot quietly inflate the false-positive rate. Refuses to pre-register vague hypotheses ("model A is better"), open-ended stopping rules ("until results stabilize"), or post-hoc rationalizations of data already observed.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

File Explorer
4 files
SKILL.md
readonly