Skip to main content
تشغيل أي مهارة في Manus
بنقرة واحدة

pre-registering-eval-study

النجوم٢
التفرعات٠
آخر تحديث١٧ يونيو ٢٠٢٦ في ١٦:٥٢

Locks the hypothesis, falsification criteria, stopping rules, and power justification of an empirical evaluation BEFORE any data is collected or observed. Use when designing an LLM eval, an A/B experiment, a clinical or observational study, an ML model comparison, or any decision that will rest on a test statistic. Produces a one-page pre-registration document committed to a versioned location prior to the first run, so that HARKing (hypothesizing after results are known), p-hacking, optional-stopping, and outcome-switching cannot quietly inflate the false-positive rate. Refuses to pre-register vague hypotheses ("model A is better"), open-ended stopping rules ("until results stabilize"), or post-hoc rationalizations of data already observed.

التثبيت

التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.

مستكشف الملفات
4 ملفات
SKILL.md
readonly