Skip to main content
在 Manus 中运行任何 Skill
一键导入

pre-registering-eval-study

星标2
分支0
更新时间2026年6月17日 16:52

Locks the hypothesis, falsification criteria, stopping rules, and power justification of an empirical evaluation BEFORE any data is collected or observed. Use when designing an LLM eval, an A/B experiment, a clinical or observational study, an ML model comparison, or any decision that will rest on a test statistic. Produces a one-page pre-registration document committed to a versioned location prior to the first run, so that HARKing (hypothesizing after results are known), p-hacking, optional-stopping, and outcome-switching cannot quietly inflate the false-positive rate. Refuses to pre-register vague hypotheses ("model A is better"), open-ended stopping rules ("until results stabilize"), or post-hoc rationalizations of data already observed.

安装

用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。

文件资源管理器
4 个文件
SKILL.md
readonly