Run flux balance analysis (FBA/FVA) on genome-scale metabolic models (E. coli core, Recon3D, AGORA2) with COBRApy; simulate single/double gene knockouts and integrate RNA-seq expression via GIMME/iMAT. Use when predicting metabolic fluxes, finding essential…
Skills in this repository
Pavel-Kravchenko/Bioinformatics - Page 3
SkillsMP has collected 213 skills from Pavel-Kravchenko/Bioinformatics. Open a skill to review its source and details.
Pavel-Kravchenko/BioinformaticsShowing 40 of 213 collected skills.
Assign molecular formulas from accurate mass/adducts and match MS/MS spectra by cosine similarity to GNPS/MassBank/HMDB. Use for LC-MS peak annotation, MSI confidence scoring, or KEGG metabolite enrichment (MSEA).
Compute alpha/beta diversity (Shannon, Simpson, Bray-Curtis, UniFrac) from a 16S/ASV feature table with scikit-bio; PCoA ordination, PERMANOVA/ANOSIM. Use for microbiome diversity or community composition questions.
Trim adapters (cutadapt), align to miRBase with Bowtie, quantify with featureCounts, run DESeq2/CPM DE testing and seed-match target prediction. Use for miRNA-seq/small RNA FASTQ processing or miRNA target prediction.
Run mixOmics PLS-DA/sPLS-DA/DIABLO to classify samples and pick stable biomarkers from paired RNA-seq/proteomics/methylation blocks. Use for supervised multi-omics classification or DIABLO biomarker discovery.
Run MOFA2 (mofapy2/muon) to fuse RNA-seq, proteomics, methylation into latent factors; decompose per-view R2, interpret weights. Use for multi-omics integration, MOFA/MOFA2 analysis, or latent factor discovery.
Test Hardy-Weinberg equilibrium, simulate Wright-Fisher drift/selection, and compute dN/dS, Tajima's D, and Fst with NumPy/SciPy. Use for neutral theory, molecular clock divergence time, selection scans, or effective population size (Ne) questions.
Train PyTorch Geometric GCN/MPNN on SMILES-derived molecular graphs to predict properties (BBBP, solubility, toxicity); compare vs Morgan-fingerprint RF. Use for GNN property prediction or SMILES-to-graph pipelines.
Compute force-field energy terms (bond/LJ/Coulomb), run energy minimization, and QC MD/homology models (RMSD, RMSF, Ramachandran) in NumPy. Use for force-field, minimization, or homology/docking validation.
Detect PPI/co-expression modules with NetworkX/python-louvain/leidenalg (Louvain, Leiden, modularity Q) and WGCNA eigengenes. Use when clustering a gene network, computing WGCNA modules, or testing DEG/pathway enrichment on network communities.
Decode Phred+33 FASTQ quality scores, compute FastQC-style per-position QC stats, and sliding-window trim reads in Python. Use when parsing FASTQ, decoding quality ASCII, or choosing Illumina/PacBio/Nanopore.
Interpolate missing time points (Newton/cubic spline), estimate derivatives, and compute AUC via trapezoidal/Simpson/curve_fit in SciPy. Use for missing qPCR points, PK dC/dt, dose-response/ROC AUC, or Michaelis-Menten/Hill fits.
Basecall ONT POD5/FAST5 signal with Dorado (fast/hac/sup, duplex, 5mC/5hmC), QC with NanoStat/NanoPlot, filter with NanoFilt, and align with Minimap2 map-ont. Use for nanopore raw-signal processing, Q-score/length read filtering, N50 computation, or a…
Build time-scaled phylogenies with TreeTime/Augur, validate clock via root-to-tip regression, interpret BEAST2 skyline plots and phylogeography. Use when dating an outbreak, estimating TMRCA, R0, or Ne(t).
Test Hardy-Weinberg equilibrium, simulate Wright-Fisher drift/selection, and compute dN/dS, Tajima's D, Fst, LD with NumPy/SciPy. Use for allele-frequency, selection, or population-structure questions.
Build and analyze protein-protein interaction (PPI) networks from STRING DB with NetworkX: compute degree/betweenness/closeness/eigenvector centrality, classify hub and bottleneck genes, test scale-free topology, and detect network communities/modules…
Design PCR/qPCR primers with primer3-py design_primers/calc_hairpin, Bio.SeqUtils Tm, and blastn specificity checks. Use when designing PCR, qPCR, cloning, or genotyping primers, or checking Tm/dimers/specificity.
Detect TATA box/Inr/DPE promoter elements, call CpG islands (O/E ratio), and build/score PWMs for TFBS scanning. Use when finding a TATA box, calling CpG islands, or scanning a sequence for transcription factor binding sites.
Compute peptide b/y ion masses, run trypsin/PMF search, quantify LFQ protein abundance (volcano plots), and calculate PTM shifts and protein inference. Use for MS/MS peptide ID, PMF search, LFQ quantification, or PTM analysis.
Run QIIME2 16S amplicon workflows: import FASTQ, DADA2 denoise to ASVs, SILVA taxonomy, alpha/beta diversity, ANCOM-BC. Use when analyzing 16S/amplicon microbiome data or .qza/.qzv pipelines.
Build QSAR classifiers from ChEMBL IC50 data with RDKit Morgan fingerprints, Random Forest, scaffold splits, and k-NN applicability domain. Use when predicting activity/pIC50 from SMILES or building structure-activity relationship models.
Parse SMILES with RDKit, compute MW/LogP/TPSA/HBD/HBA descriptors and Lipinski Ro5, build Morgan/ECFP4 and MACCS fingerprints, score Tanimoto similarity. Use for cheminformatics, drug-likeness screening, or fingerprint similarity search.
Scan DNA for promoter/regulatory elements: TATA box regex search, CpG island detection via GC%/observed-over-expected sliding windows, and PFM to PWM (log-odds) construction/scanning for TFBS. Use when locating a TSS, calling CpG islands, building a position…
Ribo-seq: cutadapt/bowtie2 adapter+rRNA removal, plastid P-site calibration, 3-nt periodicity QC, RiboCode/ribotricer ORF calling, translation efficiency. Use when user has ribosome profiling or footprint data.
Bulk RNA-seq — STAR/HISAT2/featureCounts or Salmon/kallisto quantification, TPM/DESeq2 size-factor normalization, DESeq2/pydeseq2 DE testing. Use for RNA-seq design, count matrices, or DE analysis with DESeq2, edgeR, pydeseq2.
Correct batch effects and integrate multiple scRNA-seq AnnData datasets with Harmony (harmonypy), scVI, or BBKNN; quantify mixing with LISI/ASW/kBET and transfer cell-type labels via KNN or scANVI. Use when merging samples from different…
TF-IDF normalize scATAC-seq peak matrices and run LSI/SVD, dropping depth-correlated component 1, via SnapATAC2 or Signac/Seurat. Use when clustering 10x fragments.tsv.gz, running LSI/UMAP on ATAC data, or linking peaks to genes.
Compute Gini index and replicate LFC correlation on CRISPR sgRNA count matrices; apply DESeq2-style median-ratio normalization before MAGeCK/CRISPRcleanR. Use when QC'ing a CRISPR screen count table or flagging copy-number-biased dropout.
Build AnnData from 10x MEX, compute scanpy QC (pct_counts_mt, n_genes_by_counts), MAD-filter cells, normalize/log1p, select HVGs. Use for scRNA-seq preprocessing, doublet/dying-cell filtering, or AnnData QC before clustering.
Run a full scRNA-seq analysis in scanpy on an AnnData/10x/h5ad matrix — QC filtering (pct_counts_mt, n_genes_by_counts), normalize_total/log1p, HVG selection, PCA, neighbors/UMAP, Leiden clustering, and rank_genes_groups marker detection. Use when processing…
SNP calling pipeline: Trimmomatic trim, BWA-MEM2/HISAT2 align, samtools mpileup + bcftools call, ANNOVAR annotate (dbSNP, RefGene, 1000G, ClinVar). Use for FASTQ-to-VCF pipelines or ANNOVAR variant annotation.
Analyze Visium/Xenium/MERFISH spatial transcriptomics with Squidpy/Scanpy: QC, spatial neighbor graphs, Moran's I spatially variable genes, tissue-image plots. Use for spatial autocorrelation or SVG detection.
Run t-test/Mann-Whitney/ANOVA with scipy.stats, apply Bonferroni/BH-FDR via statsmodels, compute Cohen's d and power. Use for comparing expression/counts between groups, correcting p-values, or genomics power analysis.
Parse PDB CRYST1/header for unit cell, space group, resolution, R-factors; apply symmetry operators; pick X-ray vs cryo-EM vs NMR. Use when checking structure quality, parsing CRYST1, or choosing a method.
Classify shotgun metagenome reads to species level with Kraken2, remove host reads with Bowtie2, and re-estimate abundance with Bracken. Use when doing metagenomics taxonomic profiling, Kraken2/Bracken, or host decontamination workflows.
Write pytest tests/fixtures for bio functions and GitHub Actions CI with pytest-cov, ruff, black, mypy. Use when adding tests to a bio tool, writing conftest.py fixtures, or building tests.yml/lint.yml CI workflows.
Detect TF footprints in ATAC-seq via Tn5 offset correction, insertion-profile aggregation, footprint scoring, and pybedtools intersect/slop/closest. Use when doing TF footprinting, ATAC-seq Tn5 bias correction, footprint scoring, or motif-site meta-profile…
Order scRNA-seq cells with diffusion pseudotime (scanpy sc.tl.dpt), PAGA graphs, and RNA velocity (scVelo) on spliced/unspliced counts. Use for pseudotime, root cell selection, or RNA velocity streamline plots.
Annotate a VCF's consequence/HGVS/impact with Ensembl VEP or snpEff, then join gnomAD AF, ClinVar, and dbNSFP scores. Use when annotating a VCF, running VEP/snpEff, or parsing CSQ/ANN fields.
Run GATK/bcftools BAM-to-VCF calling, parse VCF fields, decode genotypes (GT/AD/DP/GQ), hard-filter variants, test Hardy-Weinberg equilibrium. Use for SNP/indel calling, VCF/GVCF parsing, zygosity decoding, or HWE checks.