| name | scgpt |
| description | Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology. Use this skill when: (1) Producing cell embeddings from an AnnData for clustering/integration, (2) Zero-shot or fine-tuned cell-type annotation, (3) Gene-level representation for perturbation/GRN tasks.
For probabilistic single-cell models (scVI etc.), use the scvi-tools library.
|
| license | Apache-2.0 |
| origin | openai4s |
| category | biomodels |
| requirements | ["gpu"] |
| capabilities | {"network":{"mode":"none","domains":[]}} |
| metadata | {"display-name":"scGPT","third_party":[{"kind":"weights","name":"scGPT","provider":"Wang Lab (University of Toronto)","info_url":"https://github.com/bowang-lab/scGPT"}]} |
scGPT — Single-Cell Foundation Model
Prerequisites
| Requirement | Minimum | Recommended |
|---|
| Python | 3.10+ | 3.11 |
| CUDA | 12.1+ | 12.4+ |
| GPU VRAM | 16 GB | 24 GB+ |
How to run
Loading the vocabulary and checkpoint
scGPT checkpoints are raw directories (args.json, best_model.pt,
vocab.json) — not Hugging Face hub repos. Point at the directory, not an HF
repo id.
from scgpt.tokenizer.gene_tokenizer import GeneVocab
gv = GeneVocab.from_file("/path/to/scgpt-human/vocab.json")
print(len(gv))
Embedding an AnnData
import anndata as ad
from scgpt.tasks import embed_data
adata = ad.read_h5ad("dataset.h5ad")
emb = embed_data(
adata,
model_dir="/path/to/scgpt-human",
gene_col="feature_name",
use_fast_transformer=False,
)
Output format
embed_data returns an AnnData whose .obsm["X_scGPT"] is the per-cell
embedding (n_cells × emb_dim, 512 by default). Downstream: feed to
scanpy.pp.neighbors / scanpy.tl.umap.
Remote compute
Needs ≥24 GB VRAM and the released human checkpoint (~200 MB:
args.json, best_model.pt, vocab.json). Read
compute_details({provider, mode:'read'}) for an environment with
and a pre-cached checkpoint directory, then: