Skip to main content

deeptools

NGS analysis toolkit. BAM to bigWig conversion, QC (correlation, PCA, fingerprints), heatmaps/profiles (TSS, peaks), for ChIP-seq, RNA-seq, ATAC-seq visualization.

Source facts

Repository
K-Dense-AI/scientific-agent-skills
Last source activity
October 1, 2026 at 17:15
Detected SKILL.md language
English
Stars
47,404
Forks
4,286

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.

File Explorer
10 files

Showing SKILL.md

SKILL.md
Source instructions · Read-only preview
name
deeptools
description
NGS analysis toolkit. BAM to bigWig conversion, QC (correlation, PCA, fingerprints), heatmaps/profiles (TSS, peaks), for ChIP-seq, RNA-seq, ATAC-seq visualization.
license
BSD license
allowed-tools
Read Write Edit Bash
compatibility
Requires Python 3.12+ and deepTools 4.0.0 (including pysam and pyBigWig); samtools for sorting/indexing workflows. Network needed only for installation or retrieving public input data.
metadata
{"version":"2.0","last-reviewed":"2026-09-30","upstream-version":"4.0.0","skill-author":"K-Dense Inc."}
# deepTools: NGS Data Analysis Toolkit ## Overview deepTools is a comprehensive suite of Python command-line tools designed for processing and analyzing high-throughput sequencing data. Use deepTools to perform quality control, normalize data, compare samples, and generate publication-quality visualizations for ChIP-seq, RNA-seq, ATAC-seq, MNase-seq, and other NGS experiments. **Core capabilities:** - Convert BAM alignments to normalized coverage tracks (bigWig/bedGraph) - Quality control assessment (fingerprint, correlation, coverage) - Sample comparison and correlation analysis - Heatmap and profile plot generation around genomic features - Enrichment analysis and peak region visualization ## When to Use This Skill This skill should be used when: - **File conversion**: "Convert BAM to bigWig", "generate coverage tracks", "normalize ChIP-seq data" - **Quality control**: "check ChIP quality", "compare replicates", "assess sequencing depth", "QC analysis" - **Visualization**: "create heatmap around TSS", "plot ChIP signal", "visualize enrichment", "generate profile plot" - **Sample comparison**: "compare treatment vs control", "correlate samples", "PCA analysis" - **Analysis workflows**: "analyze ChIP-seq data", "RNA-seq coverage", "ATAC-seq analysis", "complete workflow" - **Working with specific file types**: BAM files, bigWig files, BED region files in genomics context ## Quick Start For users new to deepTools, start with file validation and common workflows: ### 1. Validate Input Files Before running any analysis, validate BAM, bigWig, and BED files using the validation script: ```bash python scripts/validate_files.py --bam sample1.bam sample2.bam --bed regions.bed ``` This checks local files, readable BAM/BAI/CSI structure and coordinate order, bigWig headers, and every BED row. It does not establish assembly identity or biological suitability. ### 2. Generate Workflow Template For standard analyses, use the workflow generator to create customized scripts: ```bash # List available workflows python scripts/workflow_generator.py --list # Generate ChIP-seq QC workflow python scripts/workflow_generator.py chipseq_qc -o qc_workflow.sh \ --input-bam Input.bam --chip-bams "ChIP1.bam ChIP2.bam" # Make executable and run chmod +x qc_workflow.sh ./qc_workflow.sh ``` ### 3. Most Common Operations See `assets/quick_reference.md` for frequently used commands and parameters. ## Installation ```bash uv venv --python 3.13 .venv-deeptools uv pip install --python .venv-deeptools/bin/python deepTools==4.0.0 source .venv-deeptools/bin/activate bamCoverage --version ``` Upstream recommends conda/bioconda for full dependency resolution, especially on shared HPC systems: ```bash conda create -n deeptools -c conda-forge -c bioconda python=3.13 deeptools=4.0.0 samtools ``` The 4.0.0 PyPI release includes native macOS Intel/Apple Silicon and Linux wheels. Conda availability varies by platform; the PyPI workflow above was exercised on macOS. **4.0 migration:** The five rewritten commands use Rust. `bamCoverage`, `bamCompare`, and `multiBamSummary` no longer accept `--ignoreDuplicates`; use `--samFlagExclude 1024` only after duplicate marking. `bamCompare` no longer accepts SES. Its released Rust backend accepts RPGC despite a contradictory rolling-doc note; several advertised operations are incorrect (see review). `--exactScaling` is removed from the rewritten coverage commands because scaling now uses all reads. See [references/review.md](references/review.md) for verified contracts and remaining limits. ## Core Workflows and Tool Categories Complete command sequences for ChIP-seq QC, full ChIP-seq analysis, RNA-seq coverage, and ATAC-seq analysis — plus the BAM/bigWig processing, quality control, and visualization tool categories — are in [references/core_workflows.md](references/core_workflows.md) and [references/workflows.md](references/workflows.md). Per-tool options are in [references/tools_reference.md](references/tools_reference.md). ## Normalization Methods Choosing the correct normalization is critical for valid comparisons. Consult `references/normalization_methods.md` for comprehensive guidance. **Quick selection guide:** - **ChIP-seq coverage**: Use RPGC or CPM - **ChIP-seq comparison**: Use bamCompare with log2 and readCount - **RNA-seq bins**: Use CPM - **RNA-seq genes**: Quantify with an annotation-aware gene/transcript workflow; bamCoverage RPKM scales genomic bins, not genes - **ATAC-seq**: Use CPM for the shifted-alignment template; RPGC without extension failed on shifted BAMs in 4.0.0 **Normalization methods:** - **RPGC**: 1× genome coverage (requires --effectiveGenomeSize) - **CPM**: Counts per million mapped reads - **RPKM**: Reads per kb per million (per-bin length and library-size scaling) - **BPM**: The released 4.0.0 Rust implementation reduces to CPM; do not interpret it as gene TPM - **None**: Raw counts (not recommended for comparisons) Full explanation: `references/normalization_methods.md` ## Effective Genome Sizes RPGC normalization requires effective genome size. Examples from the **4.0.0 tagged source**, not universal assembly constants (the rolling documentation differs): | Organism | Assembly | Size | Usage | |----------|----------|------|-------| | Human | GRCh38/hg38 | 2,913,022,398 | `--effectiveGenomeSize 2913022398` | | Human | T2T/CHM13CAT_v2 | 3,117,292,070 | `--effectiveGenomeSize 3117292070` | | Mouse | GRCm39/mm39 | 2,654,621,783 | `--effectiveGenomeSize 2654621783` | | Mouse | GRCm38/mm10 | 2,652,783,500 | `--effectiveGenomeSize 2652783500` | | Zebrafish | GRCz11 | 1,368,780,147 | `--effectiveGenomeSize 1368780147` | | *Drosophila* | dm6 | 142,573,017 | `--effectiveGenomeSize 142573017` | | *C. elegans* | WBcel235/ce11 | 100,286,401 | `--effectiveGenomeSize 100286401` | Verify the exact FASTA, contig set, and mapping/filter policy before choosing a value. Details: `references/effective_genome_sizes.md` ## Common Parameters Across Tools Many deepTools commands share these options: **Performance:** - `--numberOfProcessors, -p`: Use the CPU allocation allowed by your scheduler - `max` / `max/2`: Supported values for `--numberOfProcessors`; useful under schedulers because recent deepTools releases detect CPU affinity more carefully - `--region`: Process specific regions for testing (e.g., `chr1:1:1000000`) **Read Filtering:** - `--samFlagExclude 1024`: Exclude alignments already marked duplicate (0x400); does not identify duplicates - `--minMappingQuality`: Filter by alignment quality (e.g., `--minMappingQuality 10`) - `--minFragmentLength` / `--maxFragmentLength`: Fragment length bounds - `--samFlagInclude` / `--samFlagExclude`: SAM flag filtering **Read Processing:** - `--extendReads`: Extend to fragment length (ChIP-seq: YES, RNA-seq: NO) - `--centerReads`: Center at fragment midpoint for sharper signals ## Best Practices ### File Validation **Always validate files first** using `scripts/validate_files.py` to check: - File existence and readability - BAM opens with pysam, coordinate order and every alignment decode, readable BAI/CSI index - All BED rows: nonnegative, nonempty, zero-based half-open intervals; BED6 strand - bigWig opens with pyBigWig and contains indexed signal ### Analysis Strategy 1. **Start with QC**: Run correlation, coverage, and fingerprint analysis before proceeding 2. **Test on small regions**: Use `--region chr1:1:10000000` for parameter testing 3. **Document commands**: Save full command lines for reproducibility 4. **Use consistent normalization**: Apply same method across samples in comparisons 5. **Verify genome assembly**: Ensure BAM and BED files use matching genome builds ### ChIP-seq Specific - **Choose fragment handling** for ChIP-seq: use paired-end fragment lengths or a measured single-end extension; 200 bp is an illustrative fallback - **Duplicate policy**: mark duplicates upstream, then use `--samFlagExclude 1024` when the assay warrants removal; coordinate duplication alone does not prove PCR duplication - **Check enrichment first**: Run plotFingerprint before detailed analysis - **GC correction**: Only apply if significant bias detected; never use `--samFlagExclude 1024` after GC correction ### RNA-seq Specific - **Never extend reads** for RNA-seq (would span splice junctions) - **Strand-specific**: Use `--filterRNAstrand forward/reverse` for common dUTP-style stranded libraries; confirm library orientation before interpreting strand labels - **Normalization**: CPM or per-bin RPKM for coverage tracks; these are not annotation-aware gene expression estimates ### ATAC-seq Specific - **Choose the signal first**: use `alignmentSieve --ATACshift` once for shifted alignments; this alone does not create an insertion-site track - **Use only proper pairs for shifting**: `--ATACshift` is equivalent to `--shift 4 -5 5 -4` and filters to properly paired fragments - **Fragment filtering**: Set appropriate min/max fragment lengths - **Check nucleosome pattern**: inspect the unshifted library; periodicity and relative modes depend on assay/preparation, with no universal pass threshold ### Performance Optimization 1. **Use multiple processors**: `--numberOfProcessors 8` (or available cores) 2. **Increase bin size** for faster processing and smaller files 3. **Process chromosomes separately** for memory-limited systems 4. **Pre-filter BAM files** using alignmentSieve to create reusable filtered files 5. **Use bigWig over bedGraph**: Compressed and faster to process ## Troubleshooting ### Common Issues **BAM index missing:** ```bash samtools index input.bam ``` **Out of memory:** Process chromosomes individually using `--region`: ```bash bamCoverage --bam input.bam -o chr1.bw --region chr1 ``` **Slow processing:** Increase `--numberOfProcessors` and/or increase `--binSize` **bigWig files too large:** Increase bin size: `--binSize 50` or larger ### Validation Errors Run validation script to identify issues: ```bash python scripts/validate_files.py --bam *.bam --bed regions.bed ``` Common errors and solutions explained in script output. ## Reference Documentation This skill includes comprehensive reference documentation: ### references/tools_reference.md Reference for the main deepTools commands organized by category: - BAM and bigWig processing - Quality control - Visualization - Matrix operations and filtering estimates Each tool includes: - Purpose and overview - Key parameters with explanations - Usage examples - Important notes and best practices **Use this reference when:** Users ask about specific tools, parameters, or detailed usage. ### references/workflows.md Complete workflow examples for common analyses: - ChIP-seq quality control workflow - ChIP-seq complete analysis workflow - RNA-seq coverage workflow - ATAC-seq analysis workflow - Multi-sample comparison workflow - Peak region analysis workflow - Troubleshooting and performance tips **Use this reference when:** Users need complete analysis pipelines or workflow examples. ### references/normalization_methods.md Comprehensive guide to normalization methods: - Detailed explanation of each method (RPGC, CPM, RPKM, BPM, etc.) - When to use each method - Formulas and interpretation - Selection guide by experiment type - Common pitfalls and solutions - Quick reference table **Use this reference when:** Users ask about normalization, comparing samples, or which method to use. ### references/effective_genome_sizes.md Effective genome size values and usage: - Common organism values (human, mouse, fly, worm, zebrafish) - Read-length-specific values - Calculation methods - When and how to use in commands - Custom genome calculation instructions **Use this reference when:** Users need genome size for RPGC normalization or GC bias correction. ## Helper Scripts ### scripts/validate_files.py Validates BAM, bigWig, and BED files for deepTools analysis. Checks file existence, indices, and format. **Usage:** ```bash python scripts/validate_files.py --bam sample1.bam sample2.bam \ --bed peaks.bed --bigwig signal.bw ``` **When to use:** Before starting any analysis, or when troubleshooting errors. ### scripts/workflow_generator.py Generates bash templates for deepTools 4.0.0. Templates assume coordinate-sorted, indexed, duplicate-marked BAMs; QC fragment-size analysis requires paired-end data, RNA strand labels assume dUTP libraries, and TSS plots require strand-aware BED6/GTF. Review the generated script before running. RPGC workflows require an explicit `--genome-size`. **Available workflows:** - `chipseq_qc`: ChIP-seq quality control - `chipseq_analysis`: Complete ChIP-seq analysis - `rnaseq_coverage`: Strand-specific RNA-seq coverage - `atacseq`: ATAC-seq with Tn5 correction **Usage:** ```bash # List workflows python scripts/workflow_generator.py --list # Generate workflow python scripts/workflow_generator.py chipseq_qc -o qc.sh \ --input-bam Input.bam --chip-bams "ChIP1.bam ChIP2.bam" \ --threads 8 # Run generated workflow chmod +x qc.sh ./qc.sh ``` **When to use:** Users request standard workflows or need template scripts to customize. ## Assets ### assets/quick_reference.md Quick reference card with most common commands, effective genome sizes, and typical workflow pattern. **When to use:** Users need quick command examples without detailed documentation. ## Handling User Requests ### For New Users 1. Start with installation verification 2. Validate input files using `scripts/validate_files.py` 3. Recommend appropriate workflow based on experiment type 4. Generate workflow template using `scripts/workflow_generator.py` 5. Guide through customization and execution ### For Experienced Users 1. Provide specific tool commands for requested operations 2. Reference appropriate sections in `references/tools_reference.md` 3. Suggest optimizations and best practices 4. Offer troubleshooting for issues ### For Specific Tasks **"Convert BAM to bigWig":** - Use bamCoverage with appropriate normalization - Recommend RPGC or CPM based on use case - Provide effective genome size for organism - Suggest relevant parameters (extendReads, samFlagExclude, binSize) **"Check ChIP quality":** - Run full QC workflow or use plotFingerprint specifically - Explain interpretation of results - Suggest follow-up actions based on results **"Create heatmap":** - Guide through two-step process: computeMatrix → plotHeatmap - Help choose appropriate matrix mode (reference-point vs scale-regions) - Suggest visualization parameters and clustering options **"Compare samples":** - Recommend bamCompare for two-sample comparison - Suggest multiBamSummary + plotCorrelation for multiple samples - Guide normalization method selection ### Referencing Documentation When users need detailed information: - **Tool details**: Direct to specific sections in `references/tools_reference.md` - **Workflows**: Use `references/workflows.md` for complete analysis pipelines - **Normalization**: Consult `references/normalization_methods.md` for method selection - **Genome sizes**: Reference `references/effective_genome_sizes.md` ## Example Interactions **User: "I need to analyze my ChIP-seq data"** Response approach: 1. Ask about files available (BAM files, peaks, genes) 2. Validate files using validation script 3. Generate chipseq_analysis workflow template 4. Customize for their specific files and organism 5. Explain each step as script runs **User: "Which normalization should I use?"** Response approach: 1. Ask about experiment type (ChIP-seq, RNA-seq, etc.) 2. Ask about comparison goal (within-sample or between-sample) 3. Consult `references/normalization_methods.md` selection guide 4. Recommend appropriate method with justification 5. Provide command example with parameters **User: "Create a heatmap around TSS"** Response approach: 1. Verify bigWig and gene BED files available 2. Use computeMatrix with reference-point mode at TSS 3. Generate plotHeatmap with appropriate visualization parameters 4. Suggest clustering if dataset is large 5. Offer profile plot as complement ## Key Reminders - **File validation first**: Always validate input files before analysis - **Normalization matters**: Choose appropriate method for comparison type - **Extend reads carefully**: choose measured ChIP fragment handling; omit extension for spliced RNA-seq - **Respect CPU allocation**: Set `--numberOfProcessors` to allocated cores - **Test on regions**: Use `--region` for parameter testing
View on GitHub
This SKILL.md is very large, so SkillsMP previews the first section here. View on GitHub