Skip to main content

bioinformatics

Performs bioinformatics analyses including pathway enrichment, gene ontology analysis, protein-protein interaction networks, multi-omics integration, and biological sequence database querying; trigger when users discuss gene sets, biological pathways, functional annotation, or omics data integration.

跳到安装

来源信息

仓库
beita6969/ScienceClaw
最近来源活动
2026年3月12日 04:53
检测到的 SKILL.md 语言
英语
星标
907
分支
104

安装方式

默认使用会先检查来源的 Prompt;你也可以切换为直接命令,或下载本地副本。

检查来源文件

决定是否安装前,请先阅读 SKILL.md,以及 SkillsMP 当前展示的配套文件。

正在显示 SKILL.md

SKILL.md
来源说明 · 只读预览
name
bioinformatics
description
Performs bioinformatics analyses including pathway enrichment, gene ontology analysis, protein-protein interaction networks, multi-omics integration, and biological sequence database querying; trigger when users discuss gene sets, biological pathways, functional annotation, or omics data integration.
## When to Trigger Activate this skill when the user mentions: - Pathway analysis, KEGG, Reactome, WikiPathways - Gene Ontology (GO) enrichment, biological process, molecular function - Protein-protein interaction (PPI) networks, STRING, BioGRID - Multi-omics integration (transcriptomics + proteomics + metabolomics) - Gene set enrichment analysis (GSEA), over-representation analysis (ORA) - Sequence databases, UniProt, NCBI, Ensembl queries - Single-cell RNA-seq analysis, clustering, trajectory inference ## Step-by-Step Methodology 1. **Data preparation** - Standardize gene/protein identifiers (convert to Entrez, Ensembl, or UniProt IDs as needed). Remove duplicates and handle ambiguous mappings. Verify organism and genome build. 2. **Differential analysis** - For transcriptomics: DESeq2 or edgeR (count data), limma-voom (normalized). For proteomics: limma with appropriate normalization. Apply multiple testing correction (BH-FDR). Set thresholds (|log2FC| > 1, padj < 0.05 as defaults, adjustable). 3. **Functional enrichment** - Perform GO enrichment (BP, MF, CC) using clusterProfiler, g:Profiler, or DAVID. Run KEGG/Reactome pathway enrichment. Use GSEA for ranked gene lists (no arbitrary cutoff). Report enriched terms with gene ratio, p-value, adjusted p-value, and gene members. 4. **Network analysis** - Build PPI networks from STRING (confidence > 0.7 for high confidence). Identify hub genes (degree centrality), bottleneck nodes (betweenness centrality), and functional modules (MCODE, Louvain clustering). Overlay expression data on network. 5. **Multi-omics integration** - For paired omics: correlation analysis, canonical correlation (CCA), or MOFA/DIABLO. Map features across omics layers using shared identifiers or known biological connections. Identify convergent pathways. 6. **Single-cell analysis** - QC filtering (genes/cell, UMI/cell, mitochondrial %). Normalization (scran, SCTransform). Dimensionality reduction (PCA, UMAP). Clustering (Leiden, Louvain). Cell type annotation (SingleR, scType, marker genes). Trajectory inference (Monocle3, Slingshot). 7. **Visualization** - Generate volcano plots, heatmaps (with hierarchical clustering), dot plots (enrichment), network diagrams, UMAP/tSNE plots (single-cell), and circos plots (multi-omics). ## Key Databases and Tools - **Gene Ontology (GO)** - Functional annotations - **KEGG / Reactome / WikiPathways** - Pathway databases - **STRING / BioGRID / IntAct** - PPI databases - **Ensembl / NCBI / UniProt** - Sequence and annotation databases - **clusterProfiler / g:Profiler / DAVID** - Enrichment tools - **Seurat / Scanpy** - Single-cell analysis frameworks - **Cytoscape** - Network visualization ## Output Format - Enrichment results as tables: term, description, gene ratio, p-value, padj, gene list. - Volcano plots with labeled significant genes and fold-change thresholds. - Network figures with node coloring (expression), size (degree), and module highlighting. - UMAP/tSNE plots with cluster labels and cell type annotations. - Heatmaps with dendrograms and annotation bars. ## Quality Checklist - [ ] Gene ID mapping verified (conversion losses reported) - [ ] Background gene set appropriate for enrichment analysis - [ ] Multiple testing correction applied (BH-FDR or equivalent) - [ ] Redundant GO terms handled (semantic similarity, REVIGO) - [ ] Network confidence threshold specified and justified - [ ] Single-cell QC thresholds documented - [ ] Batch effects assessed and corrected if present - [ ] Results cross-validated across databases or methods - [ ] Biological interpretation grounded in literature
在 GitHub 查看