一键导入
lemma-leakage
Audit a dataset or pipeline for the five leakages that inflate a metric: target, preprocessing, temporal, group, and sampling.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Audit a dataset or pipeline for the five leakages that inflate a metric: target, preprocessing, temporal, group, and sampling.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
| name | lemma-leakage |
| description | Audit a dataset or pipeline for the five leakages that inflate a metric: target, preprocessing, temporal, group, and sampling. |
| homepage | https://github.com/tkpratardan/lemma |
| license | MIT |
Rank leakage findings by severity and show reproducible before/after scores, the contaminated path, the corrected validation design, and remaining risk.
Do not rationalize an implausible score, rely only on feature importance, randomly split temporal or repeated-entity data, or continue tuning before material leakage is resolved.
For the full leakage taxonomy, read references/deep-guide.md.
Establish a dumb baseline and an honest validation harness before any real model, so every later number means something.
Rigor for causal questions and A/B tests (the effect of acting on X): confounding, post-treatment bias, valid control groups.
Rigor for descriptive and diagnostic analytics (what happened and why): denominators, grain, and confounded slices, not model leakage.
EDA kickoff for a fresh dataset: fixed opening scaffold (goal, imports, load, sanity), then chapters derived from the data; scan leakage, land a baseline.
Rigor for statistical inference (is the difference real): hypothesis tests, power, multiple comparisons, effect size over p-value.
Final modeling once the baseline and feature set are locked: tune against validation, audit overfitting, touch the test set once, justify the complexity.