Segments the genome into chromatin states from combinatorial histone modification and chromatin factor ChIP-seq data. Uses ChromHMM (multivariate HMM on binarized signal, v1.27), Segway (Dynamic Bayesian Network on continuous signal), EpiSegMix (flexible-distribution HMM with duration modeling, 2024), EpiLogos (multi-biosample visualization), IDEAS (cell-type-aware joint), and full-stack ChromHMM (Vu Ernst 2022) for cross-cell-type segmentations. Handles state-count selection (15 vs 18 vs 25 states), binarization choice, OverlapEnrichment / NeighborhoodEnrichment downstream analysis, and cross-biosample integration. Use when learning chromatin states from a histone mark panel, characterizing learned states by genomic feature enrichment, or comparing chromatin landscapes across cell types.
Installer avec Codex ou Claude Copiez ce prompt, collez-le dans Codex, Claude ou un autre assistant, puis laissez-le vérifier la page du skill et l'installer pour vous.
Une commande directe contourne le prompt de vérification. Examinez la source avant de l'exécuter.
Segments the genome into chromatin states from combinatorial histone modification and chromatin factor ChIP-seq data. Uses ChromHMM (multivariate HMM on binarized signal, v1.27), Segway (Dynamic Bayesian Network on continuous signal), EpiSegMix (flexible-distribution HMM with duration modeling, 2024), EpiLogos (multi-biosample visualization), IDEAS (cell-type-aware joint), and full-stack ChromHMM (Vu Ernst 2022) for cross-cell-type segmentations. Handles state-count selection (15 vs 18 vs 25 states), binarization choice, OverlapEnrichment / NeighborhoodEnrichment downstream analysis, and cross-biosample integration. Use when learning chromatin states from a histone mark panel, characterizing learned states by genomic feature enrichment, or comparing chromatin landscapes across cell types.
"Integrate multiple histone modification ChIP-seq tracks into chromatin states" -> Learn a small set of recurring combinatorial patterns of histone marks (active promoter, active enhancer, poised enhancer, polycomb-repressed, heterochromatic, transcribed, etc.) and segment the genome by which state each region belongs to. Output: per-state genomic intervals, state-by-mark emission matrix, and state-state transition matrix.
Visualization across biosamples: EpiLogos (Meuleman lab)
Cell-type-aware joint: IDEAS
Chromatin state segmentation requires a panel of histone marks; minimum 4-5 marks (e.g., H3K4me3, H3K27ac, H3K4me1, H3K36me3, H3K27me3) for meaningful states. With fewer marks, simpler peak-based annotation (chipseq/peak-annotation) is more appropriate.
Practical workflow: Train at N=15, 18, 25; compare emission matrices; choose the smallest N where biology is interpretable. Higher N risks over-segmentation (state splitting random variation).
Segway Workflow
# Segway reads only the Genomedata format, so first pack the signal tracks# (one track per bigWig) plus the genome sequence into an archive.
genomedata-load \
-s hg38.fa \
-t h3k4me3=h3k4me3.bw \
-t h3k27ac=h3k27ac.bw \
-t h3k4me1=h3k4me1.bw \
-t h3k36me3=h3k36me3.bw \
-t h3k27me3=h3k27me3.bw \
signal.genomedata
# Train Segway model; GENOMEDATA and TRAINDIR are positional
segway train \
--num-labels=25 \
--num-instances=3 \
--resolution=100 \
signal.genomedata traindir/
# Posterior probabilities + hard-call annotation (GENOMEDATA TRAINDIR OUTDIR, positional)
segway posterior signal.genomedata traindir/ posteriordir/
segway annotate signal.genomedata traindir/ identifydir/
# Output: identifydir/segway.bed.gz (state assignments)
Segway uses continuous signal (from the genomedata archive) vs ChromHMM's binarized bins. Trade-off: more information per region (continuous) but more complex training.
EpiLogos Visualization
EpiLogos doesn't perform segmentation; it visualizes existing ChromHMM/Segway segmentations across many biosamples (epilogos.org).
# Use precomputed ChromHMM segmentations across multiple cell types# Web interface: https://epilogos.altius.org/# Local: github.com/meuleman/epilogos
The full-stack model trained on 1032 datasets / 127 reference epigenomes:
# Use precomputed model from Ernst lab# github.com/ernstlab/full_stack_ChromHMM_annotations# Annotate new sample by applying model to binarized data
java -mx16G -jar ChromHMM.jar MakeSegmentation \
full_stack_model_100states.txt \
binarized_sample/ \
full_stack_output/
Useful for: applying a comprehensive cross-tissue annotation to a new sample; comparing to canonical Roadmap states.
Per-Tool Failure Modes
ChromHMM -- Bin size 200 bp too coarse for sharp boundaries
Trigger: Studying TF binding boundaries or sharp enhancer transitions at 200 bp resolution.
Mechanism: ChromHMM default 200 bp bins; biology may shift within a bin.
Fix: Reduce to -b 100 or -b 50 (smaller bin); increases memory and compute time but improves boundary resolution. Re-train model at finer resolution.
ChromHMM -- Binarization throws away signal quantitation
Trigger: Distinguishing low- from high-signal regions of the same state.
Mechanism: ChromHMM binarizes each 200 bp bin to 0/1 per mark; state assignment uses combinatorial pattern, not magnitude.
Fix: Use Segway (continuous signal) or EpiSegMix (flexible distributions) for magnitude-aware segmentation.
ChromHMM / Segway -- Wrong state count
Trigger: Training with N=50 states on a 4-mark panel; or N=10 on a 7-mark panel.
Mechanism: Excess states fragment biology; insufficient states force unrelated regions into the same state.
Symptom: Emission matrix shows redundant states (multiple states with same emission profile) at high N; or biologically distinct regions lumped together at low N.
Fix: Train at N=15, 18, 25; inspect emission matrix similarity; choose the smallest N where states are interpretable as distinct biology.
Mark panel mismatch with model
Trigger: Applying Roadmap 25-state model to a sample with different mark panel.
Mechanism: Model was trained on specific marks; emission probabilities are mark-specific. Applying to different mark panel produces nonsensical state assignments.
Fix: Either train a new model on the available mark panel; OR ensure the exact same marks (and ordering) as used in the model.
IgG control instead of input
Trigger: Using IgG controls in BinarizeBam for histone marks.
Mechanism: ChromHMM's binarization compares mark signal to control; histone mark biology assumes input (sonicated chromatin) as background, not IgG.
Fix: Use sonicated input as control for histone mark ChIP. IgG is not appropriate for ChromHMM binarization of histone marks.
Cross-cell-type state mapping
Trigger: Training separate models per cell type and trying to compare state assignments.
Mechanism: State 5 in cell type A may not correspond to state 5 in cell type B if trained independently.
Fix: Train one model on concatenated data from all cell types (joint segmentation); or apply a single precomputed model (Roadmap 15-state, full-stack) to all samples for consistent state labels.
Reconciliation
Pattern
Likely cause
Action
ChromHMM and Segway segments differ
Different bin sizes / binarization vs continuous
Both can be valid; inspect emission matrices; pick the tool matching the resolution needs
State assignment varies wildly between replicates
Insufficient marks; over-binned
Increase mark panel; reduce state count
Active TSS state overlaps polycomb state at promoters
Bivalent biology (Bernstein 2006)
Expected for ESC-like cells; not an error; consider bivalent-specific state in N=18 model
Roadmap 25-state model annotates unknown cell type
Cross-cell-type generalization
Use cautiously; verify against tissue-specific tracks
Heterochromatin (H3K9me3) state has too many bins
H3K9me3 covers large fraction of genome
Expected; heterochromatin is genome-wide
Common Errors
Error / symptom
Cause
Solution
java.lang.OutOfMemoryError
Insufficient JVM heap
java -mx32G -jar ChromHMM.jar ...
BinarizeBam very slow
Large BAMs without index
samtools index all BAMs first
All states have similar emissions
Mark panel too small
Need at least 5 marks for canonical 15-state model
Segments file empty for some chromosomes
Chromosome not in chromsizes file
Add or use -chrom flag to restrict
State labels don't match Roadmap
Trained model independently
Use Roadmap precomputed model OR map states by emission similarity
ChromHMM "no signal in marks"
All bins binarized to 0
Check signal quality; verify control normalization
References
Ernst J & Kellis M 2012 Nat Methods 9:215 (ChromHMM v1)
Ernst J & Kellis M 2017 Nat Protoc 12:2478 (ChromHMM protocol)
Hoffman MM et al 2012 Nat Methods 9:473 (Segway)
Schmitz JE, Aggarwal N, Laufer L, Walter J, Salhab A, Rahmann S 2024 Bioinformatics 40:btae178 (EpiSegMix)
Meuleman W et al 2020 Nature 584:244 (EpiLogos / DHS index)
Zhang Y & Hardison 2016 Nucleic Acids Res 44:6721 (IDEAS)
Mammana A & Chung HR 2015 Genome Biol 16:151 (EpiCSeg)