一键导入
structured-intelligence
structured-intelligence 收录了来自 Scientific-Tooling 的 55 个 skills,并提供仓库级职业覆盖和站内 skill 详情页。
这个仓库中的 skills
Decision-grade enzyme/protein mutation design for thermostability with bioactivity-preserving constraints enforced by default, plus structure-aware and consensus-ranking workflows.
Create or update repository skills that conform to local templates, provider metadata requirements, registry rules, and validation workflows.
Assess data quality by reporting missing values, outliers, sample size, and variance structure for each variable in a table.
Estimate model parameters using Bayesian inference (MCMC via Stan or PyMC), returning posterior distributions and credible intervals.
Cluster samples or features using k-means or hierarchical clustering, evaluate cluster quality with silhouette scores, and produce a dendrogram or cluster plot.
Compare a continuous variable between two groups with automatic selection of t-test, Welch test, or Mann-Whitney U test based on data properties.
Fit a generalized linear model (Poisson, Binomial, or Gaussian family) for count or non-normally distributed outcome data.
Fit a linear regression model, report coefficients and model fit, and generate residual diagnostic plots.
Learn the structure and conditional probabilities of a Bayesian network from observational data for causal discovery or dependency modeling.
Embed high-dimensional data into 2D using t-SNE or UMAP for exploratory visualization of sample or cell relationships.
Compute pairwise correlations and significance tests across all numeric variables, and produce a correlation heatmap.
Perform PCA on a numeric data matrix to reduce dimensionality, visualize sample structure, and identify top contributing features.
Estimate and compare survival functions using Kaplan-Meier curves and log-rank test, with optional Cox proportional hazards regression.
Characterize the distribution of one or more numeric variables with summary statistics, normality tests, and diagnostic plots.
Compare a continuous variable across three or more groups using one-way ANOVA or Kruskal-Wallis with post-hoc testing.
Fit a logistic regression model for binary outcomes, report odds ratios and model performance (AUC-ROC, confusion matrix).
Assess alignment quality with coverage depth, mapping rates, insert size, and on-target metrics using samtools, mosdepth, and picard.
Align short reads to a reference genome with BWA-MEM2, sort and index with samtools, and mark duplicates with picard.
Annotate variants with functional effects using SnpEff or Ensembl VEP, adding gene impact, population frequencies, and clinical significance.
Call SNVs and indels from aligned BAMs using GATK HaplotypeCaller or DeepVariant with optional GVCF output for joint genotyping.
Filter raw variant calls using GATK VQSR, GATK hard filters, or bcftools expression-based filtering.
RNA-seq-specific alignment quality assessment using RSeQC and Qualimap for gene body coverage, strandedness, and rRNA contamination.
Statistical differential expression analysis using DESeq2 or edgeR in R, with support for count matrix or tximport input.
Gene set enrichment and over-representation analysis using clusterProfiler, gseapy, or fgsea, with GO, KEGG, and Reactome pathway databases.
Splice-aware alignment of RNA-seq reads to a reference genome using STAR or HISAT2.
Generate gene-level read count matrices from aligned BAMs using featureCounts or HTSeq-count.
Alignment-free transcript-level quantification using Salmon or kallisto for fast and accurate RNA-seq expression estimates.
Decision-grade protein thermal stability prediction (Tm, thermophilicity, ΔTm) from sequence and/or structure, with multi-tool evidence integration and calibrated interpretation.
Download files from the NCI Genomic Data Commons (GDC) using the GDC Data Transfer Tool and REST API, supporting both open-access and controlled-access data.
Download supplementary data files, series matrices, and raw FASTQ links from NCBI GEO for a given GSE series or GSM sample accession.
Download raw sequencing data from NCBI SRA as FASTQ files using prefetch and fasterq-dump, with support for single accessions, batch lists, and BioProject expansion.
Search the EMBL-EBI European Nucleotide Archive (ENA) for sequencing studies, runs, samples, and experiments using the ENA Portal REST API.
Search NCBI Gene Expression Omnibus (GEO) for expression datasets, series, samples, and platforms using E-utilities or the GEO query API.
Search CNCB Genome Sequence Archive (GSA) for sequencing runs, experiments, studies, and samples deposited at the China National Center for Bioinformation.
Search NCBI Datasets for genome assemblies, genes, taxonomy records, and virus genomes using the NCBI Datasets CLI or REST API.
Search NCBI Sequence Read Archive (SRA) for sequencing runs, experiments, studies, and samples using E-utilities or the NCBI Datasets CLI.
Guide users through installing and verifying bioinformatics software for NGS, RNA-seq, scRNA-seq, and metagenomics pipelines using conda, pip, or manual installation.
Marker-based and reference-based cell type annotation of single-cell RNA-seq clusters using SingleR, Azimuth, or manual curation.
Generate feature-barcode count matrices from raw scRNA-seq FASTQ files using Cell Ranger, STARsolo, or alevin-fry.
Normalization, highly variable gene selection, PCA, UMAP, and Leiden/Louvain clustering of single-cell RNA-seq data.