| name | ngs-analysis |
| description | Next-generation sequencing data analysis pipelines including bulk RNA-seq, scRNA-seq preprocessing, variant calling, and quality control. Use when working with FASTQ files, alignment (STAR, BWA), quantification (featureCounts, Salmon), DESeq2/edgeR analysis, or building NGS pipelines. Supports GEO/SRA data retrieval. |
| license | Proprietary |
NGS Data Analysis Pipelines
Data Retrieval from GEO/SRA
conda install -c bioconda sra-tools
prefetch SRR12345678
fastq-dump --split-files --gzip SRR12345678
fasterq-dump --split-files -e 8 SRR12345678
gzip SRR12345678_*.fastq
Quality Control
fastqc -t 8 -o fastqc_output/ *.fastq.gz
multiqc fastqc_output/ -o multiqc_report/
fastp -i R1.fastq.gz -I R2.fastq.gz \
-o R1_trimmed.fastq.gz -O R2_trimmed.fastq.gz \
--detect_adapter_for_pe --thread 8 \
--html fastp_report.html
Bulk RNA-seq Pipeline
Alignment with STAR
STAR --runMode genomeGenerate \
--genomeDir star_index/ \
--genomeFastaFiles genome.fa \
--sjdbGTFfile genes.gtf \
--runThreadN 16
STAR --runThreadN 16 \
--genomeDir star_index/ \
--readFilesIn R1.fastq.gz R2.fastq.gz \
--readFilesCommand zcat \
--outFileNamePrefix sample_ \
--outSAMtype BAM SortedByCoordinate \
--quantMode GeneCounts
Quantification with featureCounts
featureCounts -T 8 -p -B -C \
-a genes.gtf \
-o counts.txt \
*.bam
Salmon Pseudo-alignment
salmon index -t transcripts.fa -i salmon_index -k 31
salmon quant -i salmon_index -l A \
-1 R1.fastq.gz -2 R2.fastq.gz \
-p 8 -o salmon_quant/
Differential Expression with DESeq2
library(DESeq2)
library(tidyverse)
counts <- read.table("counts.txt", header=TRUE, row.names=1)
coldata <- read.csv("sample_info.csv", row.names=1)
dds <- DESeqDataSetFromMatrix(
countData = counts,
colData = coldata,
design = ~ condition
)
keep <- rowSums(counts(dds) >= 10) >= 3
dds <- dds[keep,]
dds <- DESeq(dds)
res <- results(dds, contrast=c
res_df as.data.frameres
rownames_to_column
filterpadj
arrangepadj
sig_genes res_df
filterpadj log2FoldChange
write.csvres_df row.names
Variant Calling (Somatic)
bwa-mem2 index reference.fa
bwa-mem2 mem -t 16 reference.fa R1.fq.gz R2.fq.gz | \
samtools sort -@ 8 -o aligned.bam
gatk MarkDuplicates -I aligned.bam -O marked.bam -M metrics.txt
gatk BaseRecalibrator -R ref.fa -I marked.bam \
--known-sites known_sites.vcf -O recal.table
gatk ApplyBQSR -R ref.fa -I marked.bam \
--bqsr-recal-file recal.table -O recal.bam
gatk Mutect2 -R ref.fa -I tumor.bam -I normal.bam \
-normal normal_sample -O somatic.vcf.gz
Python Integration
import pandas as pd
import subprocess
from pathlib import Path
def run_pipeline(fastq_dir, output_dir, genome_index):
"""Run complete RNA-seq pipeline"""
fastq_files = list(Path(fastq_dir).glob("*_R1.fastq.gz"))
for r1 in fastq_files:
sample = r1.stem.replace("_R1.fastq", "")
r2 = r1.parent / f"{sample}_R2.fastq.gz"
cmd = f"""
STAR --runThreadN 16 --genomeDir {genome_index} \
--readFilesIn {r1} {r2} --readFilesCommand zcat \
--outFileNamePrefix {output_dir}/{sample}_ \
--outSAMtype BAM SortedByCoordinate
"""
subprocess.run(cmd, shell=True, check=True)
See references/conda_envs.md for environment setup.
See scripts/batch_pipeline.py for parallel processing.