Skip to main content Skills Marketplace Discover and explore AI skills built by the community.
Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.
Copy promptShow prompt details A direct command skips the review prompt. Inspect the source before running it.
npx skills add https://github.com/mdbabumiamssm/LLMs-Universal-Life-Science-and-Clinical-Skills- --skill bio-workflows-outbreak-pipelineThe command stays on one line. Scroll horizontally to inspect it before copying.
Prefer a local copy? Download the files currently available to SkillsMP.
Download Zip Downloading... More from this repository Operate MedSAM2 for promptable segmentation of 3D medical images and medical videos, including CT lesion propagation, MRI volumes, RECIST-guided prompts, efficient CPU-oriented variants, training, and 3D Slicer integration. Use when generating or validating volumetric masks from sparse prompts or propagating masks through image slices or video frames.
Build reproducible healthcare imaging pipelines with Project MONAI for DICOM, NIfTI, pathology, and multidimensional imaging tasks including preprocessing, augmentation, training, sliding-window inference, evaluation, model bundles, labeling, and deployment. Use when implementing medical image classification, segmentation, registration, detection, generative, or foundation-model workflows in PyTorch.
Operate Google TxGemma prediction and chat models for therapeutic property prediction across small molecules, proteins, nucleic acids, diseases, targets, and cell lines. Use when formatting Therapeutics Data Commons tasks, choosing TxGemma model size or variant, running local or Model Garden inference, fine-tuning on private therapeutic data, or evaluating TxGemma in drug-discovery workflows.
Related occupations SOC
Based on SOC occupation classification
name bio-workflows-outbreak-pipeline description End-to-end outbreak investigation from pathogen isolates to transmission networks. Orchestrates MLST typing, AMR surveillance, phylodynamic dating, and transmission inference with TransPhylo. Use when investigating disease outbreaks or tracking pathogen transmission chains. tool_type mixed primary_tool mlst workflow true depends_on ["epidemiological-genomics/pathogen-typing","epidemiological-genomics/amr-surveillance","epidemiological-genomics/phylodynamics","epidemiological-genomics/transmission-inference","epidemiological-genomics/variant-surveillance"] qc_checkpoints [{"after_typing":"Valid ST assigned, cgMLST distance matrix computed"},{"after_amr":"AMR genes identified with >90% identity"},{"after_phylodynamics":"Root-to-tip R2 >0.5, clock rate plausible"},{"after_transmission":"Transmission pairs consistent with epi data"}]
Outbreak Pipeline
Complete workflow for genomic epidemiology: from pathogen isolates to transmission networks and outbreak characterization.
Workflow Overview
Pathogen Isolate Genomes (FASTA/FASTQ)
|
v
+---------+---------+
| |
v v
[1a. MLST Typing] [1b. AMR Detection] <-- Parallel execution
| |
+--------+----------+
|
v
[2. Core Genome Alignment] --> snippy / ParSNP
|
v
[3. Phylodynamics] --> TreeTime / BEAST2
|
v
[4. Transmission Inference] --> TransPhylo
|
v
Transmission Network + R0 Estimates + Timeline
Prerequisites
conda install -c bioconda mlst abricate snippy iqtree fasttree
pip install treetime transphylo biopython pandas matplotlib
Rscript -e "install.packages('TransPhylo')"
Primary Path: Bacterial Outbreak Investigation
Step 1a: MLST Typing (Parallel)
#!/bin/bash
ISOLATES="isolate1.fasta isolate2.fasta isolate3.fasta"
OUTDIR="outbreak_results"
mkdir -p ${OUTDIR} /{mlst,amr,alignment,phylo,transmission}
echo "=== MLST Typing ==="
for fasta in $ISOLATES ; do
sample=$(basename $fasta .fasta)
mlst $fasta > ${OUTDIR} /mlst/${sample} .mlst.txt
done
cat ${OUTDIR} /mlst/*.mlst.txt > ${OUTDIR} /mlst/all_mlst.tsv
echo "MLST complete: ${OUTDIR} /mlst/all_mlst.tsv"
Step 1b: AMR Detection (Parallel)
echo
fasta ;
sample=$( .fasta)
abricate --db ncbi > /amr/ .amr.tsv
abricate --summary /amr/*.amr.tsv > /amr/amr_summary.tsv
"=== AMR Detection ==="
for
in
$ISOLATES
do
basename
$fasta
$fasta
${OUTDIR}
${sample}
done
${OUTDIR}
${OUTDIR}
echo
"AMR summary: ${OUTDIR} /amr/amr_summary.tsv"
Step 2: Core Genome Alignment echo "=== Core Genome Alignment ==="
REFERENCE="reference.gbk"
for fasta in $ISOLATES ; do
sample=$(basename $fasta .fasta)
snippy --outdir ${OUTDIR} /alignment/snippy_${sample} \
--ref $REFERENCE \
--ctgs $fasta \
--cpus 8
done
snippy-core --ref $REFERENCE ${OUTDIR} /alignment/snippy_*
mv core.* ${OUTDIR} /alignment/
echo "Core alignment: ${OUTDIR} /alignment/core.aln"
Step 3: Phylodynamics with TreeTime import subprocess
from Bio import Phylo, AlignIO
import pandas as pd
import matplotlib.pyplot as plt
from pathlib import Path
outdir = Path('outbreak_results' )
subprocess.run([
'iqtree2' , '-s' , str (outdir / 'alignment/core.aln' ),
'-m' , 'GTR+G' , '-bb' , '1000' , '-nt' , 'AUTO' ,
'--prefix' , str (outdir / 'phylo/outbreak' )
], check=True )
metadata = pd.DataFrame({
'name' : ['isolate1' , 'isolate2' , 'isolate3' , 'isolate4' , 'isolate5' ],
'date' : ['2024-01-15' , '2024-01-22' , '2024-02-01' , '2024-02-10' , '2024-02-15' ]
})
metadata.to_csv(outdir / 'phylo/metadata.tsv' , sep='\t' , index=False )
subprocess.run([
'treetime' ,
'--tree' , str (outdir / 'phylo/outbreak.treefile' ),
'--aln' , str (outdir / 'alignment/core.aln' ),
'--dates' , str (outdir / 'phylo/metadata.tsv' ),
'--outdir' , str (outdir / 'phylo/treetime_output' ),
'--coalescent' , 'skyline' ,
'--clock-filter' , '3'
], check=True )
print ('TreeTime output:' , outdir / 'phylo/treetime_output' )
Step 4: Transmission Inference with TransPhylo library( TransPhylo)
library( ape)
tree <- read.nexus( "outbreak_results/phylo/treetime_output/timetree.nexus" )
dateT <- 2024.2
w_shape <- 2
w_scale <- 7/ 365
res <- inferTTree( tree, dateT = dateT,
w.shape = w_shape, w.scale = w_scale,
mcmcIterations = 10000 ,
startNeg = 1 , startPi = 0.5 )
ttree <- extractTTree( res)
medTTree <- medTTree( res)
pdf( "outbreak_results/transmission/transmission_tree.pdf" , width= 10 , height= 8 )
plotTTree( medTTree)
dev.off( )
wiw <- computeMatWIW( res)
write.csv( wiw, "outbreak_results/transmission/who_infected_whom.csv" )
R0 <- getOffspringMulti( res)
cat( "R0 estimate:" , mean( R0) , "(95% CI:" , quantile( R0, 0.025 ) , "-" , quantile( R0, 0.975 ) , ")\n" )
Python Alternative: TransPhylo via rpy2 import rpy2.robjects as ro
from rpy2.robjects.packages import importr
from rpy2.robjects import pandas2ri
import pandas as pd
from pathlib import Path
pandas2ri.activate()
transphylo = importr('TransPhylo' )
ape = importr('ape' )
outdir = Path('outbreak_results' )
tree = ape.read_nexus(str (outdir / 'phylo/treetime_output/timetree.nexus' ))
date_t = 2024.2
w_shape = 2
w_scale = 7 /365
res = transphylo.inferTTree(tree, dateT=date_t, w_shape=w_shape, w_scale=w_scale,
mcmcIterations=10000 , startNeg=1 , startPi=0.5 )
med_tree = transphylo.medTTree(res)
ro.r(f'''
pdf("{outdir} /transmission/transmission_tree.pdf", width=10, height=8)
plotTTree(medTTree({res} ))
dev.off()
''' )
print (f'Transmission tree saved to {outdir} /transmission/' )
Visualization: Outbreak Timeline import pandas as pd
import matplotlib.pyplot as plt
import matplotlib.dates as mdates
from datetime import datetime
metadata = pd.read_csv('outbreak_results/phylo/metadata.tsv' , sep='\t' )
metadata['date' ] = pd.to_datetime(metadata['date' ])
mlst = pd.read_csv('outbreak_results/mlst/all_mlst.tsv' , sep='\t' , header=None ,
names=['file' , 'scheme' , 'ST' ] + [f'locus{i} ' for i in range (7 )])
mlst['sample' ] = mlst['file' ].apply(lambda x: x.split('/' )[-1 ].replace('.fasta' , '' ))
amr = pd.read_csv('outbreak_results/amr/amr_summary.tsv' , sep='\t' )
combined = metadata.merge(mlst[['sample' , 'ST' ]], left_on='name' , right_on='sample' )
fig, ax = plt.subplots(figsize=(12 , 6 ))
colors = {'ST11' : 'red' , 'ST258' : 'blue' , 'ST307' : 'green' }
for st in combined['ST' ].unique():
subset = combined[combined['ST' ] == st]
ax.scatter(subset['date' ], [1 ]*len (subset), label=f'ST{st} ' ,
s=100 , c=colors.get(f'ST{st} ' , 'gray' ), alpha=0.7 )
ax.set_xlabel('Date' )
ax.set_ylabel('' )
ax.set_title('Outbreak Timeline by Sequence Type' )
ax.legend()
ax.xaxis.set_major_formatter(mdates.DateFormatter('%Y-%m-%d' ))
plt.xticks(rotation=45 )
plt.tight_layout()
plt.savefig('outbreak_results/outbreak_timeline.pdf' )
Parameter Recommendations Step Parameter Value Rationale snippy --mincov 10 Minimum coverage for variant call IQ-TREE -m GTR+G General time-reversible model TreeTime --clock-filter 3 Remove temporal outliers >3 IQR TransPhylo w.shape, w.scale 2, 7/365 Generation time ~7 days for many bacteria TransPhylo mcmcIterations 10000+ Ensure convergence
Troubleshooting Issue Likely Cause Solution No MLST match Novel ST or poor assembly Check assembly quality, submit novel ST Poor temporal signal Insufficient sampling, recombination Remove recombination with Gubbins, check dates TreeTime clock-filter removes many Wrong root, contamination Re-root tree, check sample quality TransPhylo non-convergence Wrong generation time Adjust w.shape/w.scale, increase iterations Missing AMR genes Database mismatch Try multiple databases (ncbi, card, resfinder)
Output Files File Description mlst/all_mlst.tsvSequence types for all isolates amr/amr_summary.tsvAMR gene presence/absence matrix alignment/core.alnCore genome SNP alignment phylo/outbreak.treefileML phylogenetic tree phylo/treetime_output/Dated tree and molecular clock transmission/transmission_tree.pdfInferred transmission network transmission/who_infected_whom.csvTransmission probability matrix
Related Skills
epidemiological-genomics/pathogen-typing - MLST and cgMLST details
epidemiological-genomics/amr-surveillance - AMRFinderPlus, ResFinder
epidemiological-genomics/phylodynamics - TreeTime, BEAST2 parameters
epidemiological-genomics/transmission-inference - TransPhylo configuration
epidemiological-genomics/variant-surveillance - Nextclade for viral outbreaks
phylogenetics/modern-tree-inference - IQ-TREE2 model selection