Skip to main content

ab-test-design-data-analysis

Designs an A/B test from scratch. Defines the hypothesis, calculates required sample size, specifies randomization method, sets success metrics and minimum detectable effect, and produces a results interpretation template. Use when the user wants to set up a controlled experiment to test whether a change produces a measurable improvement. Do NOT use for interpreting completed test results (use hypothesis-testing), analyzing existing data without experimentation (use correlation-analysis), or survey design (use survey-design in the research cluster).

跳到安装

来源信息

仓库
FerroxLabs/murage
最近来源活动
2026年9月3日 05:58
检测到的 SKILL.md 语言
英语
星标
0
分支
0

安装方式

默认使用会先检查来源的 Prompt;你也可以切换为直接命令,或下载本地副本。

检查来源文件

决定是否安装前,请先阅读 SKILL.md,以及 SkillsMP 当前展示的配套文件。

文件资源管理器
2 个文件

正在显示 SKILL.md

SKILL.md
来源说明 · 只读预览
name
ab-test-design-data-analysis
description
Designs an A/B test from scratch. Defines the hypothesis, calculates required sample size, specifies randomization method, sets success metrics and minimum detectable effect, and produces a results interpretation template. Use when the user wants to set up a controlled experiment to test whether a change produces a measurable improvement. Do NOT use for interpreting completed test results (use hypothesis-testing), analyzing existing data without experimentation (use correlation-analysis), or survey design (use survey-design in the research cluster).
license
Apache-2.0
metadata
{"author":"foundry-skills","version":"1.0.0","tags":"statistics analysis research","category":"data-analysis","subcategory":"exploratory-data-analysis","depends":"","disclaimer":"none","difficulty":"advanced"}
# A/B Test Design ## When to Use **Use this skill when:** - A user wants to measure the causal effect of a single change -- new button copy, revised pricing page, redesigned onboarding flow, different email subject line, modified recommendation algorithm -- before shipping it to all users - A user asks how many users, sessions, or impressions they need to run a valid experiment ("how long should I run this test?", "is my sample size big enough?") - A user needs a pre-registered experiment plan with a formal hypothesis, power calculation, and decision rules -- especially in regulated environments (clinical, fintech) where post-hoc analysis is unacceptable - A user wants to know whether their proposed effect size is realistic and detectable given their actual traffic volume - A user is designing a holdout test, a feature-flag rollout, or a switchback experiment and needs the same statistical rigor as a standard A/B test - A user needs to communicate an experiment plan to engineering, product, or leadership and requires a structured, defensible document - A user is computing sample sizes for a test involving continuous metrics like revenue per user, session duration, or engagement score -- not just binary conversion rates **Do NOT use when:** - The user already has test results and wants to compute a p-value or confidence interval -- use `hypothesis-testing` instead - The user wants to find patterns, correlations, or trends in existing data without running a controlled experiment -- use `correlation-analysis` - The user wants to design a survey or questionnaire to collect self-reported opinions -- use `survey-design` - The user needs to analyze test results where assignment was not properly randomized (quasi-experiments, observational studies) -- use `causal-inference` or `regression-discontinuity` if available - The user is describing a test with more than 4 simultaneously running variants -- recommend multivariate testing (MVT) or Taguchi methods instead, and note this skill covers up to 4 variants with corrections - The user wants to optimize continuously across many parameter combinations -- recommend a bandit algorithm framework (Thompson Sampling, UCB1) rather than a fixed-horizon A/B test - The user is asking about before-after comparisons without a concurrent control group -- time-series-based causal inference (difference-in-differences, interrupted time series) is more appropriate --- ## Process ### Step 1 -- Elicit the Hypothesis and Change Description Ask the user to precisely describe the intervention. Vague hypotheses produce uninterpretable results. - **What is changing?** Push for a single, atomic change. "Redesigning the checkout page" is too broad. "Changing the CTA button text from 'Submit Order' to 'Complete Purchase'" is testable. If multiple elements change simultaneously, the test cannot attribute effects to any single cause. - **What is the expected mechanism?** Ask why the change should work. "The new button text reduces cognitive friction and increases commitment language, which should increase conversion." This prevents post-hoc rationalization of any result. - **Direction of expected effect:** Establish whether the hypothesis is directional (one-tailed: "we expect improvement") or non-directional (two-tailed: "we don't know if this will help or hurt"). Default to two-tailed unless there is a strong, pre-specified reason for one-tailed. - **Write the formal null and alternative hypothesis:** - H₀: The treatment produces no change in the primary metric (metric_B - metric_A = 0) - H₁ (two-tailed): The treatment produces a nonzero change (metric_B - metric_A ≠ 0) - H₁ (one-tailed, if justified): metric_B > metric_A - **Identify the unit of analysis:** What entity experiences the change and is measured? This must match the randomization unit. If users are randomized but sessions are counted, you have a unit of analysis mismatch that inflates false positives. ### Step 2 -- Define and Operationalize the Primary Metric One primary metric. Exactly one. Pre-specify it before any data collection begins. - **Binary (proportion) metrics:** Conversion rate (converted / exposed), click-through rate (clicks / impressions), activation rate (activated / registered). These follow binomial distributions and use proportion-based sample size formulas. - **Continuous metrics:** Average order value (AOV), revenue per user (RPU), session duration, pages per session, engagement score. These follow approximately normal distributions (by CLT for large n) and use mean-comparison formulas requiring the population standard deviation. - **Count metrics:** Number of purchases per user, messages sent, items added to cart. These follow Poisson or negative binomial distributions. For count outcomes, use the delta method or bootstrap for variance estimation rather than naive t-test formulas. - **Ratio metrics:** Revenue per session (total revenue / total sessions across all users in a group). These are NOT the same as per-user averages. Ratio metrics require the delta method for variance: Var(R) ≈ (1/n) * [Var(Y) + R² * Var(X) - 2R * Cov(X,Y)] / mean(X)². Use this when the denominator varies by user. - **Specify the metric precisely:** "Conversion rate" is ambiguous. "Proportion of unique users who completed a purchase within 24 hours of landing on the product page, among users who viewed the product page at least once" is a metric. Ambiguity in metric definition leads to disputes after results are in. - **Identify secondary metrics** (3-5 maximum) that will be observed but NOT used to make the go/no-go decision. Label them exploratory. Analyzing them without correction is hypothesis generation, not confirmation. ### Step 3 -- Establish the Minimum Detectable Effect (MDE) The MDE is a business decision, not a statistical one. Drive this conversation carefully. - **Start with business impact, not statistics:** "If conversion increases from 3.0% to 3.1%, how much incremental annual revenue does that generate at your current traffic?" Walk through the math with the user. If the answer is $50,000 and the engineering cost of implementing the change is $200,000, the MDE of 0.1 pp is not worth testing for. - **Anchor to operational thresholds:** What is the minimum improvement that would cause the business to act? That is the MDE. Tests should be designed to detect effects worth caring about, not just any nonzero effect. - **Express MDE in both absolute and relative terms:** - Absolute: 1.0 percentage point (e.g., 4% → 5%) - Relative: 25% lift (1.0 / 4.0 = 25%) - For continuous metrics: minimum meaningful difference in raw units (e.g., $2.00 increase in AOV from a baseline of $45.00) - **Common MDE traps:** - Setting MDE too small to be feasible given traffic: results in a 6-month test - Setting MDE based on what the team thinks will happen: this is circular. Set MDE based on what matters, then check if the expected effect exceeds MDE - Using relative MDE for low baselines: a 20% relative lift on a 0.5% conversion rate is 0.1 pp absolute. That may require 80,000 users per group. - **For teams with no prior data:** Use industry benchmarks as a starting anchor, but always validate against the user's own historical data. Button-click optimization MDE benchmarks vary by industry from 0.5 pp to 5 pp. ### Step 4 -- Calculate Required Sample Size The sample size calculation is the technical core of the design. Get it right. **For binary (proportion) metrics:** The exact formula for two-proportions z-test (two-sided, equal groups): ``` n_per_group = (Z_α/2 + Z_β)² × [p₁(1-p₁) + p₂(1-p₂)] / (p₂ - p₁)² ``` Where: - p₁ = baseline conversion rate (control) - p₂ = p₁ + MDE (treatment, under H₁) - Z_α/2 = 1.960 for α = 0.05 (two-tailed) - Z_β = 0.842 for 80% power; 1.282 for 90% power; 0.524 for 70% power - At 80% power, α = 0.05: (1.960 + 0.842)² = 7.849 **Quick reference: sample size per group (α=0.05 two-tailed, 80% power)** | Baseline Rate | MDE (absolute) | Relative Lift | n per group | |--------------|----------------|---------------|-------------| | 1% | 0.3 pp | 30% | ~5,400 | | 2% | 0.5 pp | 25% | ~6,200 | | 3% | 0.5 pp | 17% | ~9,900 | | 3% | 1.0 pp | 33% | ~2,700 | | 5% | 0.5 pp | 10% | ~18,500 | | 5% | 1.0 pp | 20% | ~4,700 | | 10% | 1.0 pp | 10% | ~14,700 | | 10% | 2.0 pp | 20% | ~3,700 | | 20% | 2.0 pp | 10% | ~15,600 | | 20% | 4.0 pp | 20% | ~3,900 | | 50% | 5.0 pp | 10% | ~6,100 | Note: these values are precisely computed from the formula above, not approximations. The 16 × p(1-p)/MDE² shortcut underestimates sample size when p₁ and p₂ differ meaningfully. **For continuous (mean) metrics:** ``` n_per_group = 2 × (Z_α/2 + Z_β)² × σ² / δ² ``` Where: - σ = population standard deviation (estimate from historical data) - δ = minimum meaningful difference in raw units (the MDE for continuous metrics) - The factor of 2 accounts for both groups having variance σ² At 80% power, α = 0.05: n = 2 × 7.849 × σ² / δ² = 15.7 × (σ/δ)² The ratio σ/δ is the inverse of the standardized effect size (Cohen's d). Cohen's conventions: small = 0.2, medium = 0.5, large = 0.8. At d = 0.2 (small effect), n ≈ 394 per group. At d = 0.5 (medium), n ≈ 64 per group. **Practical variance estimation for continuous metrics:** - Pull 4+ weeks of historical data for the metric - Compute the per-user mean and standard deviation - Warn if the metric has high right-skew (as revenue metrics often do) -- consider using log-transformed revenue or capping outliers at the 99th percentile before analysis - For highly skewed metrics, the t-test is still valid for large n (CLT), but very large samples (n > 1,000 per group) are often needed before the normal approximation is reliable **Power settings:** - 80% power: industry default for product experiments. 20% chance of missing a real effect. - 90% power: recommended when missing an effect is costly (major platform changes, pricing tests). Increases sample size by ~35% compared to 80%. - 70% power: acceptable only for exploratory tests or when traffic is severely constrained. 30% miss rate is high. **Significance level (α) settings:** - α = 0.05: standard for product experiments - α = 0.01: use when false positives are costly (medical devices, financial products, security changes) - α = 0.10: use only for very early-stage exploratory tests where speed matters more than rigor ### Step 5 -- Design the Randomization Strategy Randomization validity determines the entire causal claim of the experiment. - **Randomization unit -- the most important decision:** - **User-level (preferred):** Randomize by persistent user ID. Prevents the same user from seeing both variants, which causes carry-over contamination. Use for logged-in product tests. - **Cookie-level:** For logged-out or anonymous users. Cookies can be cleared, leading to ~5-15% user re-assignment. Acceptable for short tests; problematic for 3+ week tests. - **Session-level:** Valid only when sessions are truly independent (e.g., each session is a separate service call). Never use session-level randomization for UI tests -- users will see flickering between variants across sessions. - **Device-level:** Intermediate between cookie and user. Stable within a device but a user on multiple devices sees both variants. Measure cross-device contamination rate if possible. - **Cluster-level (geographic or organizational):** Required when individual-level randomization is impossible (e.g., testing a policy change, testing a UI that affects how users see each other's content). Use cluster-randomized trial design -- the sample size formula changes: n_clusters = (standard n) × (1 + (m-1) × ICC), where m is cluster size and ICC is the intracluster correlation coefficient. - **Assignment mechanism -- hash-based determinism:** - Assign variant using: variant = hash(user_id + experiment_id) mod 100 - Buckets 0-49 → control, 50-99 → treatment (for 50/50 split) - The experiment_id salt ensures the same users are not always in the same group across experiments - This guarantees determinism (same user always sees same variant) without a lookup table - **Traffic allocation:** - 50/50 is statistically optimal -- it minimizes the total sample needed for a given per-group sample - Unequal splits (e.g., 90/10) are used to limit treatment exposure for risky changes. The formula for the required per-group sample at unequal allocation f (fraction to treatment): n_control = n_standard × (1 + 1/f) / (2 × f_treatment / f_control), or simply use the exact formula with unequal variances - At 90/10 split, the effective sample is dominated by the smaller group. A 90/10 test requires roughly 10× the treatment-group exposure compared to 50/50 to achieve the same power. Quantify this trade-off explicitly. - For 80/20 splits: multiply the 50/50 per-group sample by ~1.25 for the same power - **Exclusion criteria -- define before launch:** - Internal users and QA accounts (they behave atypically) - Users who have already been exposed to the new feature via a previous test or soft launch - Bots and crawlers (filter by user-agent patterns, behavioral signals) - Users in other simultaneously running experiments that affect the same surface (overlap bias) - New users vs. returning users (specify which population is in scope) - **Experiment contamination checks:** - Run an A/A test before launch if possible: assign users to two identical control groups and verify the primary metric shows no significant difference. False positive rate should be ~5%. - Check assignment balance: within 48 hours of launch, verify that the actual sample size split is close to the intended split (within ±2 percentage points). Significant imbalance suggests a logging or assignment bug. - Check covariate balance (SRM detection): if device type, country, acquisition channel, or new/returning user status is significantly imbalanced between groups, suspect a Sample Ratio Mismatch (SRM). SRM invalidates causal inference. Use a chi-squared goodness-of-fit test on the intended vs. observed allocation. ### Step 6 -- Establish Test Duration and Calendar Constraints Duration is determined by traffic, not by impatience. - **Minimum duration formula:** - Total sample needed = 2 × n_per_group (for 50/50) - Daily eligible traffic = total daily traffic × fraction eligible for this test - Minimum days = total sample needed / daily eligible traffic - Add buffer: multiply by 1.1 to 1.2 to account for traffic volatility - **Weekly cycle requirement:** - Always run for a minimum of 7 full days, even if sample size is reached in 3 days. Day-of-week effects are pervasive in consumer products: weekend users differ from weekday users in intent, device, and behavior. - Round up to the nearest full week: a test needing 10 days should run 14 days. - **Novelty effect consideration:** - New features often generate inflated engagement from curious users in the first 1-3 days. If you expect a novelty effect, extend the test to 2-3 weeks and analyze the last half of the test period separately to see if the effect stabilizes. - Compare the effect size in week 1 vs. week 2. If it decays significantly, the long-run effect is closer to week 2. - **Maximum duration:** - 4 weeks is a practical upper limit for most product experiments. Beyond 4 weeks: seasonal effects compound, the underlying user population drifts, and feature development typically moves on. - For tests exceeding 4 weeks: reframe the problem. Can the MDE be relaxed? Can a larger change be tested that would be detectable sooner? - **Calendar contamination risks to document:** - Major holidays or promotional events (Black Friday, Cyber Monday, back-to-school) - Product launches or major announcements that shift baseline traffic composition - Scheduled A/B tests on overlapping surfaces running concurrently - Infrastructure changes (server migrations, CDN changes) that could affect performance ### Step 7 -- Define Guardrail Metrics and the Decision Framework Pre-commit to the decision rules. Changing them after seeing results is p-hacking. **Guardrail metrics:** - Guardrails are metrics that must not degrade, regardless of what happens to the primary metric. Their purpose is to catch hidden negative side effects. - Select 3-5 guardrails from this taxonomy: - **Revenue guardrails:** Revenue per user, average order value, subscription cancellation rate, refund rate. Even a test designed to improve UX should not silently hurt revenue. - **Quality guardrails:** Error rate (4xx/5xx responses), page load time (P75 and P95, not just mean -- tail latency matters), crash rate for mobile features. - **Engagement guardrails:** Session depth, return visit rate, support ticket rate. A change that boosts one-time conversion but hurts retention is often a bad trade. - **User experience guardrails:** Rage clicks (rapid repeated clicking on an element), scroll depth, form abandonment rate. - Set threshold for each guardrail: "Revenue per user must not decline by more than 5% relative in the treatment group compared to control." Use this as an absolute stopping criterion. - Distinguish guardrails from secondary metrics: guardrails trigger automatic test stops; secondary metrics are observed and reported but do not trigger stops. **Pre-specified decision rules:** - Define the decision logic before the experiment starts. Write it down. Get stakeholder sign-off. - Standard two-sided decision framework: - p < α AND direction is positive → ship treatment - p < α AND direction is negative → kill treatment, investigate root cause - p ≥ α → inconclusive (see handling below) - Any guardrail breach → stop test regardless of primary metric result - For the inconclusive outcome, specify in advance: "If p ≥ 0.05 after full sample collection, we will accept no meaningful difference exists at our MDE and will NOT ship the change." Choosing to extend indefinitely after an inconclusive result is p-hacking. - If sequential testing is required (see edge cases), specify the sequential boundary or spending function before launch -- not after seeing early results. --- ## Output Format ``` ## A/B Test Design Plan ### [Test Name / Experiment ID] Prepared: [Date] --- ### 1. Hypothesis **Null hypothesis (H₀):** Changing [description of change] has no effect on [primary metric]. **Alternative hypothesis (H₁):** Changing [description of change] will [increase/decrease] [primary metric] by at least [MDE absolute] ([MDE relative]%). Test direction: [two-tailed / one-tailed]. **Mechanism:** [One sentence explaining why this change should produce the expected effect.] --- ### 2. Experiment Design | Element | Specification | |---------|--------------| | Control (A) | [Current version -- describe precisely] | | Treatment (B) | [Modified version -- describe precisely] | | Unit of randomization | [User ID / Cookie / Session / Cluster] | | Eligible population | [Definition: logged-in users who visit checkout page, etc.] | | Exclusion criteria | [Bots, internal users, users in overlapping experiments, etc.] | | Traffic split | [50/50 or specify ratio with justification] | | Assignment method | [Hash-based: hash(user_id + experiment_salt) mod 100] | --- ### 3. Metrics **Primary metric:** | Field | Value | |-------|-------| | Metric name | [Name] | | Metric type | [Binary proportion / Continuous mean / Ratio metric] | | Precise definition | [Numerator / denominator / time window] | | Baseline value | [Current observed value with confidence interval if available] | | Standard deviation | [For continuous metrics only -- from historical data] | | Minimum detectable effect | [Absolute: X pp / units] -- [Relative: X%] | **Guardrail metrics:** | Metric | Baseline | Stop-test threshold | Owner | |--------|----------|---------------------|-------| | [Metric 1] | [Value] | [≥ X% decline triggers stop] | [Team] |
在 GitHub 查看
这个 SKILL.md 很大,SkillsMP 这里只预览前一段内容。 在 GitHub 查看