| name | bio-small-rna-seq-smrna-preprocessing |
| description | Preprocess small RNA sequencing data with adapter trimming and size selection optimized for miRNA, piRNA, and other small RNAs. Use when preparing small RNA-seq reads for downstream quantification or discovery analysis. |
| tool_type | cli |
| primary_tool | cutadapt |
Small RNA Preprocessing
Adapter Trimming with Cutadapt
Small RNA libraries have specific 3' adapters that must be removed:
cutadapt \
-a TGGAATTCTCGGGTGCCAAGG \
-m 18 \
-M 30 \
--discard-untrimmed \
-o trimmed.fastq.gz \
input.fastq.gz
Common Small RNA Adapters
| Kit | 3' Adapter Sequence |
|---|
| Illumina TruSeq | TGGAATTCTCGGGTGCCAAGG |
| NEBNext | AGATCGGAAGAGCACACGTCT |
| QIAseq | AACTGTAGGCACCATCAAT |
| Lexogen | TGGAATTCTCGGGTGCCAAGGAACTCCAGTCAC |
Size Selection
cutadapt \
-a TGGAATTCTCGGGTGCCAAGG \
-m 18 -M 26 \
-o mirna_length.fastq.gz \
input.fastq.gz
Quality Trimming
cutadapt \
-q 20 \
-a TGGAATTCTCGGGTGCCAAGG \
-m 18 \
-o trimmed.fastq.gz \
input.fastq.gz
Using fastp for Small RNA
fastp \
--in1 input.fastq.gz \
--out1 trimmed.fastq.gz \
--adapter_sequence TGGAATTCTCGGGTGCCAAGG \
--length_required 18 \
--length_limit 30 \
--html report.html
Collapse Identical Reads
For small RNAs, collapsing identical sequences reduces computation:
seqkit rmdup -s trimmed.fastq.gz -o collapsed.fasta
fastx_collapser -i trimmed.fastq -o collapsed.fasta
Python Preprocessing
import gzip
from collections import Counter
def collapse_reads(fastq_path):
'''Collapse identical sequences and count occurrences'''
counts = Counter()
with gzip.open(fastq_path, 'rt') as f:
while True:
header = f.readline()
if not header:
break
seq = f.readline().strip()
f.readline()
f.readline()
if 18 <= len(seq) <= 26:
counts[seq] += 1
return counts
def write_collapsed_fasta(counts, output_path):
with open(output_path, 'w') as f:
for i, (seq, count) in enumerate(counts.most_common()):
f.write(f'>seq_{i}_x{count}\n{seq}\n')
QC Metrics for Small RNA
Key metrics to check:
- Read length distribution (should peak at 21-23 nt for miRNA)
- Adapter content (high if library is good)
- Percentage of reads in target size range
import matplotlib.pyplot as plt
from collections import Counter
def plot_length_distribution(fastq_path):
lengths = Counter()
with gzip.open(fastq_path, 'rt') as f:
for i, line in enumerate(f):
if i % 4 == 1:
lengths[len(line.strip())] += 1
plt.bar(lengths.keys(), lengths.values())
plt.xlabel('Read Length')
plt.ylabel('Count')
plt.title('Small RNA Length Distribution')
plt.savefig('length_dist.png')
Related Skills
- mirdeep2-analysis - Novel miRNA discovery
- mirge3-analysis - Fast miRNA quantification
- read-qc/adapter-trimming - General adapter trimming