Skip to main content

bio-workflows-causal-genomics-pipeline

End-to-end post-GWAS causal inference pipeline orchestrating heritability partitioning, genetic correlation, Mendelian randomization with CHP-aware sensitivity (CAUSE / LHC-MR), colocalization, fine-mapping with SuSiE / FOCUS, mediation, TWAS triangulation, cis-pQTL drug-target MR, effector-gene prioritization (L2G / PoPS / cS2G), and GenomicSEM common-factor GWAS. Use when triangulating causal inference across multiple complementary methods, prioritizing tissues via stratified LDSC, nominating or de-risking drug targets, mapping a lead SNP to a candidate effector gene, modeling shared genetic architecture across correlated traits, or producing a STROBE-MR-compliant publication-grade evidence battery from GWAS summary statistics.

ソース情報

リポジトリ
GPTomics/bioSkills
ソースの最終更新活動
2026年7月16日 20:43
検出された SKILL.md の言語
英語
スター
1,209
フォーク
251

インストール方法

デフォルトでは、最初にソースを確認する Prompt が選択されています。直接コマンドに切り替えるか、ローカルコピーをダウンロードすることもできます。

ソースファイルを確認

インストールを決める前に、SKILL.md と SkillsMP に表示されている付属ファイルをお読みください。

ファイルエクスプローラー
3 ファイル

SKILL.md を表示中

SKILL.md
ソースの指示 · 読み取り専用プレビュー
name
bio-workflows-causal-genomics-pipeline
description
End-to-end post-GWAS causal inference pipeline orchestrating heritability partitioning, genetic correlation, Mendelian randomization with CHP-aware sensitivity (CAUSE / LHC-MR), colocalization, fine-mapping with SuSiE / FOCUS, mediation, TWAS triangulation, cis-pQTL drug-target MR, effector-gene prioritization (L2G / PoPS / cS2G), and GenomicSEM common-factor GWAS. Use when triangulating causal inference across multiple complementary methods, prioritizing tissues via stratified LDSC, nominating or de-risking drug targets, mapping a lead SNP to a candidate effector gene, modeling shared genetic architecture across correlated traits, or producing a STROBE-MR-compliant publication-grade evidence battery from GWAS summary statistics.
tool_type
r
primary_tool
TwoSampleMR
workflow
true
depends_on
["causal-genomics/mendelian-randomization","causal-genomics/colocalization-analysis","causal-genomics/fine-mapping","causal-genomics/pleiotropy-detection","causal-genomics/mediation-analysis","causal-genomics/transcriptome-wide-association","causal-genomics/heritability-partitioning","causal-genomics/proteome-mr-drug-target","causal-genomics/effector-gene-prioritization","causal-genomics/genetic-correlation","causal-genomics/genomic-sem"]
qc_checkpoints
[{"after_h2_ldsc":"Mean chi-squared > 1.02; h2 SE < 0.02; intercept ratio < 0.3"},{"after_rg_check":"If abs(rg) > 0.3 then CHP suspected and CAUSE/LHC-MR required"},{"after_instrument_selection":"F-statistic > 10 (two-sample) or > 20 (one-sample); no palindromic SNPs at MAF near 0.5"},{"after_mr":"IVW + Egger + weighted median + weighted mode concordance"},{"after_sensitivity":"MR-PRESSO global p, Egger intercept p (with Isq >= 0.9 for NOME), Steiger directionality"},{"after_chp_check":"CAUSE delta_elpd z > 1.96 OR LHC-MR posterior excludes null"},{"after_coloc":"PP.H4 >= 0.7 triangulation, >= 0.8 publication, >= 0.95 industry; p12 sensitivity stable"},{"after_finemapping":"Credible-set purity (min_abs_corr) >= 0.5; estimate_s_rss lambda < 0.05 if external LD"},{"after_twas":"FOCUS PIP >= 0.8 for candidate causal gene; tissue selected via stratified LDSC"},{"after_effector_gene":"L2G + PoPS + coloc + TWAS concordance >= 3 of 6 evidence streams"},{"after_mediation":"rho_crit > 0.3 OR mediational E-value > 2 (Imai sensitivity)"}]
## Version Compatibility Reference examples tested with: TwoSampleMR 0.5+, MR-PRESSO 1.0+, coloc 5.2+, susieR 0.12+, MendelianRandomization 0.9+, ldsc 1.0.1 (python3 fork), MetaXcan 0.7+, pyfocus 0.6+, MAGMA 1.10+, MRlap 0.0.3+, cause 1.2+, lhcMR 0.0.0.9000+, HDL 1.4+, LAVA 0.1+, GenomicSEM 0.0.5+. Before using code patterns, verify installed versions match. If versions differ: - R: `packageVersion('<pkg>')` then `?function_name` to verify parameters - Python: `pip show <pkg>` then `help(module.function)` to check signatures - CLI: `<tool> --version` then `<tool> --help` to confirm flags If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying. # Causal Genomics Pipeline **"Run post-GWAS causal inference from summary statistics"** -> Orchestrate heritability partitioning and tissue prioritization, genetic-correlation diagnostics, instrument selection, Mendelian randomization with CHP-aware sensitivity, colocalization, fine-mapping (SuSiE / FOCUS), mediation, TWAS triangulation, cis-pQTL drug-target MR, effector-gene prioritization, and (optionally) GenomicSEM common-factor GWAS to triangulate causal evidence and nominate publication-grade causal exposures, genes, and mechanisms. This is a workflow skill: it owns the chaining decisions and hand-offs, not the internals of any one step. Every step below cross-references the component skill that teaches its mechanism. ## The governing principle A causal claim is decided at the seams where summary statistics meet, not inside any single method. 1. **Everything shares ONE genome build, and effect alleles are harmonized once.** The exposure and outcome sumstats, the LD reference, and any eQTL/pQTL panel must share a build; if they differ, liftover ONCE, strand-aware (the BBIS inverted-region danger), BEFORE harmonization. `harmonise_data` aligns effect alleles; a flipped palindrome (MAF near 0.5) flips the causal-effect SIGN silently — drop intermediate-frequency palindromes. 2. **The LD reference ancestry must match the GWAS ancestry, committed once and inherited by clumping, coloc.susie, fine-mapping (SuSiE-rss), and TWAS.** A mismatched LD panel corrupts credible sets and colocalization silently; the in-pipeline alarms are `estimate_s_rss` lambda <0.05 and credible-set purity `min_abs_corr` >=0.5. 3. **Order is causal: a diagnostic at step 0 decides whether a later step runs.** Cross-trait LDSC `abs(rg)>0.3` makes correlated horizontal pleiotropy (CHP) suspected, so CAUSE/LHC-MR becomes MANDATORY — MR-PRESSO is BLIND to CHP. Tissue picked at step 0 (stratified LDSC) flows into TWAS; picking it post-hoc is circular. Steiger pre-filter directionality and LD-clump BEFORE coloc/fine-mapping. 4. **Triangulate; refuse the single number.** No single MR estimator, coloc PP.H4, or evidence stream is decisive — report IVW+Egger+median+mode concordance, coloc PP.H4 with a p12 sweep, and effector-gene evidence across >=3 of 6 streams. A bare headline statistic hides the seam. ## Pipeline Overview ``` GWAS Summary Statistics (exposure + outcome) | v [0. Pre-flight: h2 + tissue prioritization + rg diagnostic] LDSC / S-LDSC baseline-LD / Finucane 2018 cell-type Cross-trait LDSC / HDL / LAVA --> if abs(rg) > 0.3 then CHP-aware MR required | v [1. Instrument Selection] -----> LD clumping, F-stat filtering, Steiger pre-filter | v [2. Mendelian Randomization] --> IVW, MR-Egger, Weighted Median/Mode, MR-RAPS | +--> [3. Sensitivity] -------> MR-PRESSO, Egger intercept (Isq), leave-one-out, Steiger | +--> [3b. CHP-aware MR] -----> CAUSE (delta_elpd), LHC-MR posterior (if rg > 0.3) | v [4. Colocalization] -----------> coloc.abf / coloc.susie / HyPrColoc / SMR-HEIDI | v [5. Fine-Mapping] -------------> SuSiE rss + estimate_s_rss / FINEMAP-inf / PolyFun | v [6. Mediation Analysis] -------> Network MR / MVMR / CMAverse 4-way | +--> [7. TWAS triangulation] -> FUSION / S-PrediXcan / FOCUS PIP >= 0.8 | +--> [8. Cis-pQTL drug-target MR] -> UKB-PPP / deCODE + cross-platform replication | v [9. Effector-gene prioritization] -> Open Targets L2G + PoPS + cS2G + coloc + TWAS (>= 3 of 6 evidence) | v [10. (optional) GenomicSEM common-factor GWAS] -> factor model + Q_SNP | v Triangulated causal-evidence summary across methods ``` ## Step 0: Pre-flight - Heritability, Tissue Prioritization, Genetic Correlation ```bash ldsc.py --h2 trait.sumstats.gz --ref-ld-chr eur_w_ld_chr/ --w-ld-chr eur_w_ld_chr/ --out trait.h2 ldsc.py --h2-cts trait.sumstats.gz --ref-ld-chr baselineLD. --ref-ld-chr-cts Multi_tissue_gene_expr.ldcts --w-ld-chr weights. --out trait.cts ldsc.py --rg trait1.sumstats.gz,trait2.sumstats.gz --ref-ld-chr eur_w_ld_chr/ --w-ld-chr eur_w_ld_chr/ --out rg ``` **Goal:** Confirm heritable signal, pick the right tissue for TWAS / V2G, and detect shared heritable confounding that mandates CHP-aware MR (CAUSE / LHC-MR). **Reconciliation:** S-LDSC mean chi-squared > 1.02 with h2 SE < 0.02 and intercept ratio < 0.3 is required. Cell-type prioritization with coefficient_p < 0.05 / N_tissues nominates the tissue for downstream TWAS weights and ABC enhancer-gene priors. If cross-trait LDSC abs(rg) > 0.3 (and HDL sample-overlap < 5%), Step 3b becomes mandatory. See causal-genomics/heritability-partitioning and causal-genomics/genetic-correlation. ## Step 1: Instrument Selection ```r library(TwoSampleMR) exposure_dat <- read_exposure_data(filename = 'exposure_gwas.tsv', sep = '\t', snp_col = 'SNP', beta_col = 'BETA', se_col = 'SE', effect_allele_col = 'A1', other_allele_col = 'A2', eaf_col = 'EAF', pval_col = 'P') exposure_dat <- subset(exposure_dat, pval.exposure < 5e-8) exposure_dat <- clump_data(exposure_dat, clump_r2 = 0.001, clump_kb = 10000) exposure_dat$F_stat <- (exposure_dat$beta.exposure / exposure_dat$se.exposure)^2 exposure_dat <- subset(exposure_dat, F_stat >= 10) exposure_dat <- subset(exposure_dat, !(eaf.exposure > 0.42 & eaf.exposure < 0.58 & substr(effect_allele.exposure,1,1) %in% c('A','T') & substr(other_allele.exposure,1,1) %in% c('A','T'))) ``` For cis-MR (drug target) use clump_r2 = 0.1 within +/- 500 kb of the gene. Use 5e-9 if M > 5M variants tested. ## Step 2: Mendelian Randomization ```r outcome_dat <- read_outcome_data(filename = 'outcome_gwas.tsv', sep = '\t', snp_col = 'SNP', beta_col = 'BETA', se_col = 'SE', effect_allele_col = 'A1', other_allele_col = 'A2', eaf_col = 'EAF', pval_col = 'P') dat <- harmonise_data(exposure_dat, outcome_dat) mr_results <- mr(dat, method_list = c('mr_ivw', 'mr_egger_regression', 'mr_weighted_median', 'mr_weighted_mode')) ``` Concordance across IVW, Egger, weighted median, and weighted mode is the headline causal claim. See causal-genomics/mendelian-randomization. ## Step 3: Sensitivity Analysis ```r library(MRPRESSO) presso <- mr_presso(BetaOutcome = 'beta.outcome', BetaExposure = 'beta.exposure', SdOutcome = 'se.outcome', SdExposure = 'se.exposure', OUTLIERtest = TRUE, DISTORTIONtest = TRUE, data = dat, NbDistribution = 5000, SignifThreshold = 0.05) egger_int <- mr_pleiotropy_test(dat) isq <- Isq(abs(dat$beta.exposure), dat$se.exposure) # I^2_GX needs same-sign effects; pass abs(beta) het <- mr_heterogeneity(dat) loo <- mr_leaveoneout(dat) steiger <- directionality_test(dat) ``` Isq >= 0.9 is required for the MR-Egger NOME assumption; below that, run SIMEX correction or drop Egger. See causal-genomics/pleiotropy-detection. ## Step 3b: CHP-aware MR (when rg > 0.3 or shared confounder suspected) ```r library(cause) library(lhcMR) cause_fit <- cause(X = cause_dat, variants = top_vars, param_ests = params) elpd <- summary(cause_fit)$elpd # lhcMR is a 3-call chain: merge_sumstats (takes LD/rho paths) -> calculate_SP -> lhc_mr lhc_df <- merge_sumstats(input.files, trait.names, LD.filepath = ld_path, rho.filepath = rho_path) SP_list <- calculate_SP(lhc_df, trait.names, nStep = 2, SP_single = 3, SP_pair = 50) lhc_fit <- lhc_mr(SP_list, trait.names, paral_method = 'lapply', nBlock = 200) ``` CAUSE reports delta_elpd of sharing-vs-causal model; z > 1.96 favors true causation over CHP. LHC-MR jointly estimates causal effect and confounder effect via likelihood; the 95% credible interval excluding zero is the causal-effect verdict. Required when LDSC abs(rg) > 0.3. See causal-genomics/pleiotropy-detection. ## Step 4: Colocalization ```r library(coloc) d1 <- list(beta = exposure_locus$BETA, varbeta = exposure_locus$SE^2, snp = exposure_locus$SNP, position = exposure_locus$BP, type = 'quant', N = exposure_n, MAF = exposure_locus$EAF) d2 <- list(beta = outcome_locus$BETA, varbeta = outcome_locus$SE^2, snp = outcome_locus$SNP, position = outcome_locus$BP, type = 'cc', N = outcome_n, s = case_fraction, MAF = outcome_locus$EAF) result <- coloc.abf(d1, d2, p1 = 1e-4, p2 = 1e-4, p12 = 1e-5) sens <- coloc::sensitivity(result, rule = 'H4 > 0.7') ``` PP.H4 >= 0.7 for triangulation, >= 0.8 for publication, >= 0.95 for industry-grade target packages. Always sweep p12 over 1e-6 to 5e-5; conclusions must be stable across the sweep. For allelic heterogeneity use coloc.susie. See causal-genomics/colocalization-analysis. ## Step 5: Fine-Mapping with SuSiE ```r library(susieR) R <- as.matrix(read.csv('ld_matrix.csv', row.names = 1)) diag_s <- estimate_s_rss(z = locus_stats$BETA / locus_stats$SE, R = R, n = sample_size) fitted <- susie_rss(bhat = locus_stats$BETA, shat = locus_stats$SE, R = R, n = sample_size, L = 10, coverage = 0.95, min_abs_corr = 0.5) cs <- fitted$sets$cs ``` If `estimate_s_rss` lambda > 0.05 the external LD is mismatched; rerun with in-sample LD or use SuSiE-inf / FINEMAP-inf. min_abs_corr >= 0.5 (r-squared >= 0.25) is the purity threshold for retaining a credible set. For HLA use L = 20-30. See causal-genomics/fine-mapping. ## Step 6: Mediation Analysis ```r library(TwoSampleMR) mv_exposures <- mv_extract_exposures(c('ieu-a-2', 'ieu-a-1089')) mv_outcome <- extract_outcome_data(mv_exposures$SNP, 'ieu-a-7') mvdat <- mv_harmonise_data(mv_exposures, mv_outcome) mvmr_result <- mv_multiple(mvdat) ``` Indirect effect = total - direct. For molecular mediators (expression, methylation, protein), prefer two-step MR with cis-instruments at the mediator. Run Imai sensitivity (rho_crit) or mediational E-value > 2. See causal-genomics/mediation-analysis. ## Step 7: TWAS Triangulation ```bash python MetaXcan/SPrediXcan.py --model_db_path mashr_Whole_Blood.db \ --covariance mashr_Whole_Blood.txt.gz --gwas_file gwas.txt.gz \ --snp_column SNP --effect_allele_column A1 --non_effect_allele_column A2 \ --beta_column BETA --se_column SE --pvalue_column P \ --output_file twas.csv focus finemap gwas.sumstats.gz 1000G.EUR.QC.1 mashr.db --chr 1 --p-threshold 5e-8 --out twas.focus ``` **Goal:** Nominate gene-level causal hits and prune LD-induced TWAS false positives. Tissue is picked from Step 0 stratified LDSC; Bonferroni at 0.05 / N_tissues. FOCUS PIP >= 0.8 retains a single candidate causal gene per region; without FOCUS, co-regulated TWAS hits cannot be distinguished. Cross-reference TWAS hits with coloc.susie PP.H4 and cis-eQTL MR for triangulation. See causal-genomics/transcriptome-wide-association. ## Step 8: Cis-pQTL Drug-Target MR ```r library(TwoSampleMR) pqtl_dat <- extract_instruments(outcomes = 'prot-a-XXX', p1 = 5e-8, clump = TRUE, r2 = 0.1, kb = 1000) pqtl_dat <- subset(pqtl_dat, chr.exposure == target_chr & pos.exposure > target_tss - 500000 & pos.exposure < target_tss + 500000) # extract_instruments returns chr.exposure/pos.exposure out_dat <- extract_outcome_data(snps = pqtl_dat$SNP, outcomes = c('ieu-a-7', 'ieu-b-31')) dat <- harmonise_data(pqtl_dat, out_dat) mr_results <- mr(dat, method_list = c('mr_ivw', 'mr_wald_ratio')) ``` **Goal:** Mimic pharmacological inhibition of a drug target via cis-pQTL and triangulate with coloc. Cross-platform replication on Olink (UKB-PPP) and SomaScan (deCODE) is mandatory; the two platforms are concordant for ~60% of proteins and discordant calls are platform artifacts. Run pheWAS for on-target adverse effects, PAV-excluded sensitivity, and coloc.susie PP.H4 >= 0.8 at the cis-pQTL locus. See causal-genomics/proteome-mr-drug-target. ## Step 9: Effector-Gene Prioritization ```bash magma --bfile g1000_eur --pval gwas.tsv N=N --gene-annot genes.annot --out trait python munge_feature_directory.py --gene_annot_path genes.txt --feature_dir features/ --save_prefix pops python pops.py --gene_annot_path genes.txt --feature_mat_prefix pops --num_feature_chunks 2 --magma_prefix trait --out_prefix trait.pops ``` **Goal:** Map each fine-mapped credible set to a candidate effector gene by integrating six evidence streams.
GitHubで見る
この SKILL.md は非常に大きいため、SkillsMP では最初のセクションだけを表示しています。 GitHubで見る