| name | bio-workflows-expression-to-pathways |
| description | Workflow from differential expression results to functional enrichment analysis. Covers GO, KEGG, Reactome enrichment with clusterProfiler and visualization. Use when taking DE results to pathway enrichment. |
| tool_type | r |
| primary_tool | clusterProfiler |
| workflow | true |
| depends_on | ["pathway-analysis/go-enrichment","pathway-analysis/kegg-pathways","pathway-analysis/reactome-pathways","pathway-analysis/gsea","pathway-analysis/enrichment-visualization"] |
| qc_checkpoints | [{"input_validation":"Valid gene IDs, sufficient DE genes"},{"enrichment_qc":"Reasonable number of terms, p-values not all significant"}] |
Version Compatibility
Reference examples tested with: DESeq2 1.42+, R stats (base), ReactomePA 1.46+, clusterProfiler 4.10+, ggplot2 3.5+
Before using code patterns, verify installed versions match. If versions differ:
- R:
packageVersion('<pkg>') then ?function_name to verify parameters
If code throws ImportError, AttributeError, or TypeError, introspect the installed
package and adapt the example to match the actual API rather than retrying.
Expression to Pathways Workflow
"Find enriched pathways from my differential expression results" → Orchestrate GO enrichment (clusterProfiler), GSEA, KEGG/Reactome pathway mapping, and enrichment visualization from DE gene lists or ranked gene lists.
Convert differential expression results into biological insights through functional enrichment analysis.
Method Selection
| Scenario | Method | Why |
|---|
| Have DE results with Wald stat / t-stat for all genes | GSEA (Step 4) | Uses full ranking; no arbitrary cutoff; ~35% higher F1 than ORA |
| Clear gene list from non-DE source (co-expression, GWAS) | ORA (Steps 1-3) | No ranking available |
| RNA-seq with known gene length bias | GOseq (goseq package) | Standard ORA ignores length bias |
| Bacterial / prokaryotic data | KEGG with locus tags | No org.*.eg.db; use keyType='kegg' |
| Multiple conditions to compare | compareCluster or mitch | Never compare p-values across separate enrichments |
When in doubt, run both ORA and GSEA and compare. Concordant results are more trustworthy.
Workflow Overview
DE Results (gene list or ranked list)
|
v
[1. Gene ID Conversion] --> Convert to Entrez/Ensembl
|
v
[2. Over-representation Analysis]
|
+---> GO Enrichment (BP, MF, CC)
|
+---> KEGG Pathways
|
+---> Reactome Pathways
|
v
[3. GSEA (ranked genes)]
|
v
[4. Visualization] -----> Dot plots, networks, bar plots
|
v
Functional annotations and pathway insights
Input Preparation
From DESeq2 Results
library(DESeq2)
library(clusterProfiler)
library(org.Hs.eg.db)
res <- read.csv('deseq2_results.csv', row.names = 1)
sig_genes <- rownames(subset(res, padj < 0.05 & abs(log2FoldChange) > 1))
background_genes <- rownames(res[!is.na(res$pvalue), ])
ranked_genes <- res$stat
names(ranked_genes) <- rownames(res
ranked_genes sortranked_genesranked_genes decreasing
Gene ID Conversion
sig_entrez <- bitr(sig_genes, fromType = 'SYMBOL', toType = 'ENTREZID',
OrgDb = org.Hs.eg.db)
ranked_entrez <- bitr(names(ranked_genes), fromType = 'SYMBOL', toType = 'ENTREZID',
OrgDb = org.Hs.eg.db)
ranked_list <- ranked_genes[ranked_entrez$SYMBOL]
names(ranked_list) <- ranked_entrez$ENTREZID
Step 1: GO Over-representation Analysis
bg_entrez <- bitr(background_genes, fromType = 'SYMBOL', toType = 'ENTREZID', OrgDb = org.Hs.eg.db)
go_bp <- enrichGO(gene = sig_entrez$ENTREZID,
universe = bg_entrez$ENTREZID,
OrgDb = org.Hs.eg.db,
ont = 'BP',
pAdjustMethod = 'BH',
pvalueCutoff = 0.05,
qvalueCutoff = 0.1,
readable = TRUE)
go_mf <- enrichGO(gene = sig_entrez$ENTREZID,
universe = bg_entrez$ENTREZID,
OrgDb = org.Hs.eg.db,
ont =
pAdjustMethod
pvalueCutoff
readable
go_cc enrichGOgene sig_entrezENTREZID
universe bg_entrezENTREZID
OrgDb org.Hs.eg.db
ont
pAdjustMethod
pvalueCutoff
readable
go_bp_simple simplifygo_bp cutoff by
Step 2: KEGG Pathway Enrichment
kegg <- enrichKEGG(gene = sig_entrez$ENTREZID,
organism = 'hsa',
pvalueCutoff = 0.05,
qvalueCutoff = 0.1)
kegg <- setReadable(kegg, OrgDb = org.Hs.eg.db, keyType = 'ENTREZID')
Step 3: Reactome Pathway Enrichment
library(ReactomePA)
reactome <- enrichPathway(gene = sig_entrez$ENTREZID,
organism = 'human',
pvalueCutoff = 0.05,
readable = TRUE)
Step 4: Gene Set Enrichment Analysis (GSEA)
gsea_go <- gseGO(geneList = ranked_list,
OrgDb = org.Hs.eg.db,
ont = 'BP',
minGSSize = 10,
maxGSSize = 500,
pvalueCutoff = 0.05,
verbose = FALSE)
gsea_kegg <- gseKEGG(geneList = ranked_list,
organism = 'hsa',
minGSSize = 10,
maxGSSize = 500,
pvalueCutoff = 0.05,
verbose = FALSE)
Step 5: Visualization
library(enrichplot)
library(ggplot2)
dotplot(go_bp_simple, showCategory = 20) +
ggtitle('GO Biological Process Enrichment')
ggsave('go_bp_dotplot.pdf', width = 10, height = 8)
barplot(kegg, showCategory = 15) +
ggtitle('KEGG Pathway Enrichment')
ggsave('kegg_barplot.pdf', width = 9, height = 6)
go_bp_simple <- pairwise_termsim(go_bp_simple)
emapplot(go_bp_simple, showCategory = 30) +
ggtitle('GO Term Similarity Network'
ggsave width height
cnetplotgo_bp showCategory categorySize
ggtitle
ggsave width height
gseaplot2gsea_kegg geneSetID pvalue_table
ggsave width height
ridgeplotgsea_go showCategory
ggsave width height
Step 6: Export Results
write.csv(as.data.frame(go_bp), 'go_bp_enrichment.csv', row.names = FALSE)
write.csv(as.data.frame(kegg), 'kegg_enrichment.csv', row.names = FALSE)
write.csv(as.data.frame(reactome), 'reactome_enrichment.csv', row.names = FALSE)
write.csv(as.data.frame(gsea_go), 'gsea_go_results.csv', row.names = FALSE)
combined <- rbind(
data.frame(Database = 'GO_BP', as.data.frame(go_bp_simple)[1:10,]),
data.frame(Database as.data.framekegg
data.frameDatabase as.data.framereactome
write.csvcombined row.names
Parameter Recommendations
| Analysis | Parameter | Value |
|---|
| enrichGO | pvalueCutoff | 0.05 |
| enrichGO | qvalueCutoff | 0.1 |
| simplify | cutoff | 0.7 |
| gseGO | minGSSize | 10 |
| gseGO | maxGSSize | 500 |
| GSEA | perm | 1000 (default) |
Troubleshooting
| Issue | Likely Cause | Solution |
|---|
| No enriched terms | Too few genes, wrong IDs | Check gene IDs, relax thresholds |
| All terms significant | Too many genes | Be more stringent with DE cutoffs |
| Gene ID conversion fails | Wrong organism, format | Check OrgDb package, gene format |
| GSEA no results | Poor ranking, small gene sets | Check ranked list, adjust minGSSize |
Complete Workflow Script
library(clusterProfiler)
library(org.Hs.eg.db)
library(ReactomePA)
library(enrichplot)
library(ggplot2)
de_file <- 'deseq2_results.csv'
output_dir <- 'pathway_analysis'
dir.create(output_dir, showWarnings = FALSE)
res <- read.csv(de_file, row.names = 1)
sig_genes <- rownames(subset(res, padj < 0.05 & abs(log2FoldChange) > 1))
cat('Significant genes:', length(sig_genes), '\n')
sig_entrez <- bitr(sig_genes, fromType toType OrgDb org.Hs.eg.db
cat nrowsig_entrez
ranked resstat
ranked rownamesres
ranked sortrankedranked decreasing
ranked_entrez bitrranked fromType toType OrgDb org.Hs.eg.db
ranked_list rankedranked_entrezSYMBOL
ranked_list ranked_entrezENTREZID
bg_entrez bitrrownamesresrespvalue fromType toType OrgDb org.Hs.eg.db
go_bp enrichGOsig_entrezENTREZID universe bg_entrezENTREZID OrgDb org.Hs.eg.db ont readable
go_bp_simple simplifygo_bp cutoff
kegg enrichKEGGsig_entrezENTREZID organism
kegg setReadablekegg OrgDb org.Hs.eg.db keyType
reactome enrichPathwaysig_entrezENTREZID organism readable
gsea_go gseGOranked_list OrgDb org.Hs.eg.db ont verbose
pdffile.pathoutput_dir width height
printdotplotgo_bp_simple showCategory ggtitle
printbarplotkegg showCategory ggtitle
nrowas.data.framereactome
printdotplotreactome showCategory ggtitle
dev.off
write.csvas.data.framego_bp_simple file.pathoutput_dir row.names
write.csvas.data.framekegg file.pathoutput_dir row.names
write.csvas.data.framereactome file.pathoutput_dir row.names
cat output_dir
cat nrowas.data.framego_bp_simple
cat nrowas.data.framekegg
cat nrowas.data.framereactome
Prokaryotic Organisms
For bacteria/archaea, standard org.db annotation packages are unavailable. Use KEGG directly with strain-specific organism codes:
search_kegg_organism('Pseudomonas aeruginosa', by = 'scientific_name')
kegg_bac <- enrichKEGG(gene = sig_gene_ids, organism = 'pae', keyType = 'kegg',
pvalueCutoff = 0.05)
kegg_ko <- enrichKEGG(gene = ko_ids, organism = 'ko', keyType = 'kegg')
GO enrichment for prokaryotes: use enricher() with custom GO-to-gene mapping from eggNOG-mapper or InterProScan output, rather than org.db packages.
Multi-Condition Enrichment Comparison
When comparing enrichment across conditions (e.g., treatment A vs B vs C):
gene_clusters <- list(
ConditionA = sig_genes_A,
ConditionB = sig_genes_B,
ConditionC = sig_genes_C
)
cc <- compareCluster(gene_clusters, fun = 'enrichKEGG', organism = 'hsa')
dotplot(cc, showCategory = 10) + theme(axis.text.x = element_text(angle = 45, hjust = 1))
Do not compare raw -log10(p-values) across conditions — they scale with sample size. Compare NES (normalized enrichment scores) for GSEA, or use compareCluster for ORA.
Related Skills
- database-access/biomart-queries - Bulk SYMBOL -> Entrez / HGNC / UniProt ID mapping
- database-access/uniprot-access - UniProt ID-mapping (alternative to bitr) for obsolete-accession resolution
- database-access/interaction-databases - Layer STRING / OmniPath interactions on enriched gene sets
- database-access/ortholog-inference - Pull pre-computed orthologs for cross-species enrichment
- pathway-analysis/go-enrichment - GO enrichment details
- pathway-analysis/kegg-pathways - KEGG analysis
- pathway-analysis/reactome-pathways - Reactome analysis
- pathway-analysis/gsea - GSEA methods
- pathway-analysis/enrichment-visualization - Visualization options
- differential-expression/de-results - Preparing gene lists for enrichment