| name | bio-ribo-seq-ribosome-periodicity |
| description | Validate Ribo-seq data quality by checking 3-nucleotide periodicity and calculating P-site offsets. Use when assessing library quality or determining read offsets for downstream analysis. |
| tool_type | python |
| primary_tool | Plastid |
Ribosome Periodicity Analysis
3-Nucleotide Periodicity
Ribosomes move 3 nucleotides per codon. Good Ribo-seq data shows strong periodicity:
from plastid import BAMGenomeArray, FivePrimeMapFactory, GenomicSegment
import numpy as np
import matplotlib.pyplot as plt
alignments = BAMGenomeArray('riboseq.bam', mapping=FivePrimeMapFactory())
Calculate P-site Offset
from plastid import metagene_analysis
def determine_psite_offset(bam_path, annotation_file):
'''Determine optimal P-site offset from metagene analysis'''
from plastid import GTF2_TranscriptAssembler, BAMGenomeArray
transcripts = list(GTF2_TranscriptAssembler(annotation_file))
alignments = BAMGenomeArray(bam_path, mapping=FivePrimeMapFactory())
metagene_data = metagene_analysis(
transcripts,
alignments,
upstream=50,
downstream=100
)
return metagene_data
Metagene Plots
def plot_metagene(metagene_data, offset=12):
'''Plot metagene profile around start codon'''
fig, axes = plt.subplots(1, 2, figsize=(12, 5))
positions = np.arange(-50, 100)
for frame in range(3):
frame_positions = positions[positions % 3 == frame]
counts = metagene_data[positions % 3 == frame]
axes[0].bar(frame_positions, counts, alpha=0.7, label=f'Frame {frame}')
axes[0].set_xlabel('Position relative to start codon')
axes[0].set_ylabel('Normalized counts')
axes[0].legend()
axes[0].axvline(0, color='red', linestyle='--', label='Start')
from scipy.fft import fft
fft_result = np.abs(fft(metagene_data))
freq = np.fft.fftfreq(len(metagene_data))
axes[1].plot(1/freq[1:len(freq)//2], fft_result[1:len(freq)//2])
axes[1].set_xlabel('Period (nt)')
axes[].set_ylabel()
axes[].axvline(, color=, linestyle=)
plt.tight_layout()
plt.savefig()
Assess by Read Length
def periodicity_by_length(bam_path, annotation_file):
'''Calculate periodicity score for each read length'''
import pysam
reads_by_length = {}
with pysam.AlignmentFile(bam_path, 'rb') as bam:
for read in bam:
if not read.is_unmapped:
length = read.query_length
if length not in reads_by_length:
reads_by_length[length] = []
reads_by_length[length].append(read)
results = {}
for length, reads in reads_by_length.items():
if len(reads) > 1000:
periodicity = calculate_periodicity(reads, annotation_file)
results[length] = periodicity
return results
P-site Offset Table
Common P-site offsets by read length (5' end mapping):
| Read Length | P-site Offset |
|---|
| 28 nt | 12 |
| 29 nt | 12 |
| 30 nt | 13 |
| 31 nt | 13 |
| 32 nt | 14 |
Validate with RiboCode
RiboCode_onestep \
-g annotation.gtf \
-r riboseq.bam \
-f genome.fa \
-o output_dir
Related Skills
- riboseq-preprocessing - Generate aligned BAM
- orf-detection - Uses P-site offsets
- translation-efficiency - Requires proper positioning