| name | bio-flow-cytometry-differential-analysis |
| description | Differential abundance and state analysis for cytometry data. Compare cell populations between conditions using statistical methods. Use when testing for significant changes in cell frequencies or marker expression between groups. |
| tool_type | r |
| primary_tool | CATALYST |
Version Compatibility
Reference examples tested with: R stats (base), edgeR 4.0+, ggplot2 3.5+, limma 3.58+
Before using code patterns, verify installed versions match. If versions differ:
- R:
packageVersion('<pkg>') then ?function_name to verify parameters
If code throws ImportError, AttributeError, or TypeError, introspect the installed
package and adapt the example to match the actual API rather than retrying.
Differential Analysis
"Compare cell populations between my conditions" → Test for significant changes in cell type frequencies (differential abundance) or marker expression levels (differential state) between experimental groups.
- R:
CATALYST::testDA_edgeR() or diffcyt::testDA_GLMM()
Differential Abundance (DA)
Goal: Test which cell population clusters differ in frequency between experimental conditions.
Approach: Create a design matrix and contrast from sample metadata, then run edgeR-based differential abundance testing on cluster counts per sample using testDA_edgeR from the diffcyt framework.
library(CATALYST)
library(diffcyt)
sce <- readRDS('sce_clustered.rds')
design <- createDesignMatrix(ei(sce), cols_design = 'condition')
contrast <- createContrast(c(0, 1))
res_DA <- testDA_edgeR(sce, design, contrast, cluster_id = 'meta20')
rowData(res_DA)$cluster_id
rowData(res_DA)$p_adj
sig_DA <- rowData(res_DA)$p_adj < 0.05
table(sig_DA)
Differential State (DS)
res_DS <- testDS_limma(sce, design, contrast,
cluster_id = 'meta20',
markers_include = rownames(sce)[rowData(sce)$marker_class == 'state'])
ds_results <- rowData(res_DS)
Visualization
plotDiffHeatmap(sce, res_DA, all = TRUE, fdr = 0.05)
plotDiffHeatmap(sce, res_DS, all = TRUE, fdr = 0.05)
plotAbundances(sce, k = 'meta20', by = 'cluster_id', group_by = 'condition')
Manual Statistical Testing
library(tidyverse)
freqs <- colData(sce) %>%
as.data.frame() %>%
group_by(sample_id, condition, cluster_id = cluster_ids(sce, 'meta20')) %>%
summarise(n = n(), .groups = 'drop') %>%
group_by(sample_id) %>%
mutate(freq = n / sum(n) * 100)
test_abundance <- function(df, cluster) {
cluster_data <- filter(df, cluster_id == cluster)
ctrl <- filter(cluster_data condition freq
treat filtercluster_data condition freq
ctrl treat
test t.testtreat ctrl
data.frame
cluster cluster
fc meantreat meanctrl
pvalue testp.value
results map_dfruniquefreqscluster_id test_abundancefreqs .x
resultspadj p.adjustresultspvalue method
Mixed Effects Models
library(lme4)
library(lmerTest)
fit_mixed <- function(df, cluster) {
cluster_data <- filter(df, cluster_id == cluster)
model <- lmer(freq ~ condition + (1|patient_id), data = cluster_data)
coef <- summary(model)$coefficients
return(data.frame(
cluster = cluster,
estimate = coef[2, 'Estimate'],
pvalue = coef[2, 'Pr(>|t|)']
))
}
CITRUS (Automated Discovery)
library(citrus)
fcs_files <- list.files('data', pattern = '\\.fcs$', full.names = TRUE)
labels <- c(rep('Control', 2), rep('Treatment', 2))
citrus_result <- citrus(
fcs_files,
labels,
fileSampleSize = 1000,
featureType = 'abundances',
modelType = 'glmnet',
family = 'classification'
)
citrus_plot(citrus_result)
Volcano Plot
library(ggplot2)
da_df <- as.data.frame(rowData(res_DA))
da_df$significant <- da_df$p_adj < 0.05
ggplot(da_df, aes(x = logFC, y = -log10(p_adj), color = significant)) +
geom_point() +
geom_hline(yintercept = -log10(0.05), linetype = 'dashed') +
geom_vline(xintercept = c(-1, 1), linetype = 'dashed') +
scale_color_manual(values
theme_bw
labstitle
Export Results
da_results <- as.data.frame(rowData(res_DA))
da_results$analysis <- 'DA'
ds_results <- as.data.frame(rowData(res_DS))
ds_results$analysis <- 'DS'
write.csv(da_results, 'da_results.csv', row.names = FALSE)
write.csv(ds_results, 'ds_results.csv', row.names = FALSE)
Multiple Comparisons
design_full <- model.matrix(~ 0 + condition, data = ei(sce))
colnames(design_full) <- levels(factor(ei(sce)$condition))
contrasts <- makeContrasts(
TreatA_vs_Ctrl = TreatmentA - Control,
TreatB_vs_Ctrl = TreatmentB - Control,
TreatA_vs_B = TreatmentA - TreatmentB,
levels = design_full
)
res_list <- lapply(1:ncol(contrasts), function(i) {
testDA_edgeR(sce, design_full, contrasts[, i], cluster_id
Related Skills
- clustering-phenotyping - Cluster data first
- gating-analysis - Compare gated populations
- differential-expression/de-results - Similar statistical concepts