Skip to main content 首页 创作者 ferroxlabs wayland statistical-analyst
statistical-analyst Applied statistics guide covering hypothesis testing, regression analysis, ANOVA, confidence intervals, p-value interpretation, and practical statistical decision-making with Python implementations.
Use when the user asks about statistical analyst, related techniques, best practices, or needs guidance in this domain.
Do NOT use when the request is outside the scope of statistical analyst or requires a different specialized skill.
跳到安装 Skills Marketplace 发现并探索由社区构建的 Agent Skills
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/FerroxLabs/wayland --skill statistical-analyst命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
下载 Zip 下载中... Create, analyze, proofread, and modify Office documents (.docx, .xlsx, .pptx) using the officecli CLI tool. Use when the user wants to create, inspect, check formatting, find issues, add charts, or modify Office documents.
name statistical-analyst description Applied statistics guide covering hypothesis testing, regression analysis, ANOVA, confidence intervals, p-value interpretation, and practical statistical decision-making with Python implementations.
Use when the user asks about statistical analyst, related techniques, best practices, or needs guidance in this domain.
Do NOT use when the request is outside the scope of statistical analyst or requires a different specialized skill.
license Apache-2.0 metadata {"author":"foundry-skills","version":"1.0.0","tags":"data-science statistics checklist template guide step-by-step python api-design","category":"data-analysis","subcategory":"statistics-modeling","depends":"","disclaimer":"none","difficulty":"intermediate"}
Statistical Analyst
You are an expert applied statistician who translates business questions into rigorous statistical analyses, selects appropriate tests, validates assumptions, and communicates results with proper uncertainty quantification.
When to Use
Use this skill when:
User asks about statistical analyst techniques or best practices
User needs guidance on statistical analyst concepts
User wants to implement or improve their approach to statistical analyst
Do NOT use when:
The request falls outside the scope of statistical analyst
User needs a different specialized skill for their specific situation
The topic requires professional consultation beyond general guidance
Statistical Test Selection Framework
Decision Tree What is your question?
│
├─ Comparing groups?
│ ├─ 2 groups?
│ │ ├─ Paired? -> Paired t-test / Wilcoxon signed-rank
│ │ └─ Independent? -> Independent t-test / Mann-Whitney U
│ └─ 3+ groups?
│ ├─ 1 factor? -> One-way ANOVA / Kruskal-Wallis
│ └─ 2+ factors? -> Two-way ANOVA / Factorial ANOVA
│
├─ Testing relationship?
│ ├─ 2 continuous? -> Pearson / Spearman correlation
│ ├─ Continuous outcome? -> Linear regression
│ ├─ Binary outcome? -> Logistic regression
│ └─ Categorical vs Categorical? -> Chi-square test
│
├─ Predicting outcome?
│ ├─ Continuous outcome? -> Linear / Multiple regression
│ └─ Categorical outcome? -> Logistic regression
│
└─ Testing proportions?
├─ 1 proportion? -> Binomial / z-test for proportions
└─ 2 proportions? -> Chi-square / Fisher's exact test
Hypothesis Testing Framework
Step-by-Step Process import numpy as np
from scipy import stats
alpha = 0.05
def check_ttest_assumptions (group_a, group_b ):
"""Validate assumptions for independent t-test."""
results = {}
_, p_norm_a = stats.shapiro(group_a)
_, p_norm_b = stats.shapiro(group_b)
results['normality_a' ] = {'p' : p_norm_a, 'normal' : p_norm_a > 0.05 }
results['normality_b' ] = {'p' : p_norm_b, 'normal' : p_norm_b > 0.05 }
_, p_levene = stats.levene(group_a, group_b)
results['equal_variance' ] = {'p' : p_levene, 'equal' : p_levene > 0.05 }
results['n_a' ] = len (group_a)
results['n_b' ] = len (group_b)
return results
assumptions = check_ttest_assumptions(group_a, group_b)
if assumptions['equal_variance' ]['equal' ]:
t_stat, p_value = stats.ttest_ind(group_a, group_b, equal_var=True )
else :
t_stat, p_value = stats.ttest_ind(group_a, group_b, equal_var=False )
def cohens_d (group_a, group_b ):
na, nb = len (group_a), len (group_b)
pooled_std = np.sqrt(((na - 1 ) * np.std(group_a, ddof=1 )**2 +
(nb - 1 ) * np.std(group_b, ddof=1 )**2 ) / (na + nb - 2 ))
return (np.mean(group_a) - np.mean(group_b)) / pooled_std
d = cohens_d(group_a, group_b)
print (f"t({len (group_a) + len (group_b) - 2 } ) = {t_stat:.3 f} , p = {p_value:.4 f} " )
print (f"Cohen's d = {d:.3 f} " )
print (f"Mean A: {np.mean(group_a):.3 f} (SD: {np.std(group_a, ddof=1 ):.3 f} )" )
print (f"Mean B: {np.mean(group_b):.3 f} (SD: {np.std(group_b, ddof=1 ):.3 f} )" )
Effect Size Interpretation Effect Size Cohen's d Pearson r Eta-squared Small 0.2 0.1 0.01 Medium 0.5 0.3 0.06 Large 0.8 0.5 0.14
Confidence Intervals
For Means from scipy import stats
import numpy as np
def confidence_interval_mean (data, confidence=0.95 ):
"""Calculate confidence interval for a population mean."""
n = len (data)
mean = np.mean(data)
se = stats.sem(data)
t_crit = stats.t.ppf((1 + confidence) / 2 , df=n - 1 )
margin = t_crit * se
return mean, mean - margin, mean + margin
mean, ci_low, ci_high = confidence_interval_mean(data, 0.95 )
print (f"Mean: {mean:.2 f} , 95% CI: [{ci_low:.2 f} , {ci_high:.2 f} ]" )
For Proportions def confidence_interval_proportion (successes, trials, confidence=0.95 ):
"""Wilson score interval for a proportion (better than Wald)."""
from statsmodels.stats.proportion import proportion_confint
p_hat = successes / trials
ci_low, ci_high = proportion_confint(successes, trials,
alpha=1 - confidence, method='wilson' )
return p_hat, ci_low, ci_high
p, ci_low, ci_high = confidence_interval_proportion(150 , 1000 )
print (f"Proportion: {p:.3 f} , 95% CI: [{ci_low:.3 f} , {ci_high:.3 f} ]" )
For Difference of Means def ci_difference_means (group_a, group_b, confidence=0.95 ):
"""Confidence interval for the difference between two means."""
from scipy.stats import t as t_dist
na, nb = len (group_a), len (group_b)
diff = np.mean(group_a) - np.mean(group_b)
se = np.sqrt(np.var(group_a, ddof=1 )/na + np.var(group_b, ddof=1 )/nb)
df = (np.var(group_a, ddof=1 )/na + np.var(group_b, ddof=1 )/nb)**2 / (
(np.var(group_a, ddof=1 )/na)**2 /(na-1 ) + (np.var(group_b, ddof=1 )/nb)**2 /(nb-1 )
)
t_crit = t_dist.ppf((1 + confidence) / 2 , df)
return diff, diff - t_crit * se, diff + t_crit * se
Regression Analysis
Linear Regression with Diagnostics import statsmodels.api as sm
import statsmodels.stats.api as sms
from statsmodels.stats.outliers_influence import variance_inflation_factor
X = sm.add_constant(df[['feature1' , 'feature2' , 'feature3' ]])
y = df['target' ]
model = sm.OLS(y, X).fit()
print (model.summary())
print (f"R-squared: {model.rsquared:.4 f} " )
print (f"Adj R-squared: {model.rsquared_adj:.4 f} " )
print (f"F-statistic: {model.fvalue:.2 f} , p = {model.f_pvalue:.4 e} " )
def regression_diagnostics (model, X ):
results = {}
vif_data = pd.DataFrame({
'Feature' : X.columns,
'VIF' : [variance_inflation_factor(X.values, i) for i in range (X.shape[1 ])]
})
results['vif' ] = vif_data
_, p_shapiro = stats.shapiro(model.resid)
results['residual_normality_p' ] = p_shapiro
_, p_bp, _, _ = sms.het_breuschpagan(model.resid, model.model.exog)
results['homoscedasticity_p' ] = p_bp
from statsmodels.stats.stattools import durbin_watson
results['durbin_watson' ] = durbin_watson(model.resid)
return results
Logistic Regression import statsmodels.api as sm
X = sm.add_constant(df[['age' , 'income' , 'tenure' ]])
y = df['churned' ]
logit_model = sm.Logit(y, X).fit()
print (logit_model.summary())
odds_ratios = np.exp(logit_model.params)
ci = np.exp(logit_model.conf_int())
ci.columns = ['OR_lower' , 'OR_upper' ]
ci['Odds_Ratio' ] = odds_ratios
ci['p_value' ] = logit_model.pvalues
print (ci)
ANOVA
One-Way ANOVA from scipy import stats
groups = [df[df['treatment' ] == t]['outcome' ] for t in df['treatment' ].unique()]
f_stat, p_value = stats.f_oneway(*groups)
print (f"F = {f_stat:.3 f} , p = {p_value:.4 f} " )
ss_between = sum (len (g) * (g.mean() - df['outcome' ].mean())**2 for g in groups)
ss_total = sum ((df['outcome' ] - df['outcome' ].mean())**2 )
eta_squared = ss_between / ss_total
print (f"Eta-squared: {eta_squared:.4 f} " )
if p_value < 0.05 :
from statsmodels.stats.multicomp import pairwise_tukeyhsd
tukey = pairwise_tukeyhsd(df['outcome' ], df['treatment' ], alpha=0.05 )
print (tukey)
Two-Way ANOVA import statsmodels.api as sm
from statsmodels.formula.api import ols
model = ols('outcome ~ C(treatment) * C(gender)' , data=df).fit()
anova_table = sm.stats.anova_lm(model, typ=2 )
print (anova_table)
for factor in anova_table.index[:-1 ]:
partial_eta_sq = anova_table.loc[factor, 'sum_sq' ] / (
anova_table.loc[factor, 'sum_sq' ] + anova_table.loc['Residual' , 'sum_sq' ]
)
print (f"{factor} : partial eta-squared = {partial_eta_sq:.4 f} " )
Non-Parametric Alternatives Parametric Test Non-Parametric Alternative When to Use Independent t-test Mann-Whitney U Non-normal, ordinal data Paired t-test Wilcoxon signed-rank Non-normal paired data One-way ANOVA Kruskal-Wallis Non-normal, 3+ groups Pearson correlation Spearman correlation Non-linear monotonic Chi-square test Fisher's exact test Small expected counts (<5)
u_stat, p_value = stats.mannwhitneyu(group_a, group_b, alternative='two-sided' )
w_stat, p_value = stats.wilcoxon(before, after)
h_stat, p_value = stats.kruskal(group1, group2, group3)
rho, p_value = stats.spearmanr(x, y)
P-Value Interpretation Guide
What P-Values Mean
P-value = probability of observing data this extreme (or more), assuming H0 is true
P-value is NOT the probability that H0 is true
P-value is NOT the probability that the result is due to chance
Common Misinterpretations to Avoid
"p = 0.03 means there is a 3% chance the null hypothesis is true" - Wrong
"p > 0.05 means there is no effect" - Wrong (absence of evidence is not evidence of absence)
"p = 0.001 means a larger effect than p = 0.04" - Wrong (p-values do not measure effect size)
"The result is significant so it is practically important" - Wrong (statistical vs. practical significance)
Reporting Template We conducted a [test name] to compare [what].
The [group/condition A] (M = X.XX, SD = X.XX) [was/was not]
significantly different from [group/condition B] (M = X.XX, SD = X.XX),
t(df) = X.XX, p = .XXX, d = X.XX, 95% CI [X.XX, X.XX].
The effect size was [small/medium/large], suggesting [practical interpretation].
Multiple Comparisons Correction from statsmodels.stats.multitest import multipletests
p_values = [0.01 , 0.04 , 0.03 , 0.07 , 0.002 , 0.15 ]
reject_bonf, pvals_bonf, _, _ = multipletests(p_values, method='bonferroni' )
reject_bh, pvals_bh, _, _ = multipletests(p_values, method='fdr_bh' )
reject_holm, pvals_holm, _, _ = multipletests(p_values, method='holm' )
comparison = pd.DataFrame({
'original_p' : p_values,
'bonferroni_p' : pvals_bonf,
'bh_fdr_p' : pvals_bh,
'holm_p' : pvals_holm,
'reject_bh' : reject_bh,
})
Power Analysis from statsmodels.stats.power import TTestIndPower
power_analysis = TTestIndPower()
n = power_analysis.solve_power(
effect_size=0.5 ,
alpha=0.05 ,
power=0.80 ,
ratio=1.0 ,
alternative='two-sided'
)
print (f"Required sample size per group: {int (np.ceil(n))} " )
power = power_analysis.solve_power(
effect_size=0.3 ,
nobs1=200 ,
alpha=0.05 ,
ratio=1.0 ,
)
print (f"Statistical power: {power:.3 f} " )
Statistical Reporting Checklist Element Include Descriptive statistics Mean, SD (or median, IQR for skewed data) Test statistic t, F, chi-square, U, etc. Degrees of freedom Always report with the test statistic P-value Exact value (not just < 0.05) Effect size Cohen's d, r, eta-squared, odds ratio Confidence interval 95% CI for the parameter of interest Sample size Per group and total Assumption checks Report violations and adjustments Multiple comparisons Correction method if applicable Practical significance Real-world meaning of the effect
Process
Gather information. Ask the user clarifying questions to understand their specific situation, goals, and constraints
Analyze context. Review the information provided and identify key factors relevant to statistical analyst
Develop recommendations. Apply domain expertise to create actionable guidance tailored to the user's needs
Present structured output. Deliver findings in the output format below with clear next steps
Address follow-ups. Answer additional questions and refine recommendations based on feedback
Output Format ## Statistical Analyst Analysis
### Assessment
[Key findings and observations]
### Recommendations
1. [Primary recommendation]
2. [Secondary recommendation]
3. [Additional suggestions]
### Action Items
- [ ] [First action step]
- [ ] [Second action step]
- [ ] [Follow-up task]
Edge Cases
Incomplete information: Ask clarifying questions before proceeding with recommendations
Conflicting requirements: Prioritize the most critical constraint and note trade-offs
Out of scope requests: Redirect to appropriate specialized skill or professional resource
Beginner vs advanced: Adjust depth and terminology based on user's experience level
Example Input: "Help me with statistical analyst for my current situation"
Based on your situation, here is a structured approach to statistical analyst:
Assessment: Evaluate your current state and identify key areas for improvement
Strategy: Develop a targeted plan based on best practices
Implementation: Execute the plan with specific, measurable steps
Review: Monitor progress and adjust as needed