Skip to main content

bio-clinical-databases-hla-typing

Calls HLA class I and class II alleles at 2/4/6/8-field resolution from WGS/WES/RNA-seq/long-read data using OptiType, HLA-LA, T1K, Polysolver, HLA-HD, arcasHLA, StarPhase, or HIBAG imputation. Use when typing for HSCT, solid-organ transplant, neoantigen prediction, PGx screening (B*57:01, B*15:02, etc.), or disease-association studies, with reconciliation across tools and IPD-IMGT/HLA version mismatch handling.

来源信息

仓库
GPTomics/bioSkills
最近来源活动
2026年7月22日 12:17
检测到的 SKILL.md 语言
英语
星标
1,209
分支
251

安装方式

默认使用会先检查来源的 Prompt;你也可以切换为直接命令,或下载本地副本。

检查来源文件

决定是否安装前,请先阅读 SKILL.md,以及 SkillsMP 当前展示的配套文件。

文件资源管理器
4 个文件

正在显示 SKILL.md

SKILL.md
来源说明 · 只读预览
name
bio-clinical-databases-hla-typing
description
Calls HLA class I and class II alleles at 2/4/6/8-field resolution from WGS/WES/RNA-seq/long-read data using OptiType, HLA-LA, T1K, Polysolver, HLA-HD, arcasHLA, StarPhase, or HIBAG imputation. Use when typing for HSCT, solid-organ transplant, neoantigen prediction, PGx screening (B*57:01, B*15:02, etc.), or disease-association studies, with reconciliation across tools and IPD-IMGT/HLA version mismatch handling.
tool_type
cli
primary_tool
T1K
## Version Compatibility Reference examples tested with: OptiType 1.3.5, HLA-LA 1.0.4, T1K 1.0.6 (Song 2023), Polysolver 4.0, HLA-HD 1.7.1, arcasHLA 0.6.0, StarPhase 1.0+ (PacBio), HIBAG 1.40+, samtools 1.19+, bwa-mem 0.7.17+. IPD-IMGT/HLA database release frequency is quarterly; tools must be re-bundled with the current release to capture new alleles (~38,000 alleles at Jan 2024; ~43,000+ by Jul 2025). Before using code patterns, verify installed versions match. If versions differ: - Python: `pip show <package>` then `help(module.function)` to check signatures - CLI: `<tool> --version` then `<tool> --help` to confirm flags If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying. Tool reference-bundle vintage matters more than algorithm choice for non-European cohorts; a 2022-bundled HLA-LA will silently miss thousands of post-2022 alleles dominant in African and South Asian ancestry. # HLA Typing for Clinical Applications **'Determine HLA genotype for HSCT / neoantigen prediction / PGx screening'** -> Call HLA class I (A, B, C) and class II (DRB1, DRB3/4/5, DQA1, DQB1, DPA1, DPB1) alleles at the resolution required by the downstream application. - CLI (general-purpose all-rounder): `t1k --preset hla -1 R1.fq -2 R2.fq -f hla_reference.fa` - CLI (class I gold standard from WES/WGS): `OptiTypePipeline.py -i R1.fq R2.fq -d` - CLI (class I + II with PRG): `HLA-LA.pl --BAM input.bam --graph PRG_MHC_GRCh38_withIMGT` - CLI (RNA-seq): `arcasHLA extract sample.bam -o out && arcasHLA genotype out/sample.extracted.fq.gz` - CLI (long-read transplant-grade): PacBio HiFi StarPhase - R (imputation from SNP arrays): `HIBAG::predict()` with ancestry-stratified reference panel ## Resolution Levels and What Each Application Requires HLA nomenclature: **HLA-A\*02:01:01:01** = family : protein-changing : synonymous : intronic/UTR. **Expression suffixes:** `N` (null; DNA present, no protein expressed); `L` (low expression); `S` (secreted); `Q` (questionable); `A` (aberrant). A serologically apparent DR4-positive donor carrying `DRB4*01:03:01:02N` is functionally DR53-negative; a classic HSCT donor-selection failure. | Application | Min resolution | Why | |-------------|---------------|-----| | **HSCT (unrelated donor)** | 6-field (12/12 match) | Null alleles + permissive DPB1 + Bw4/Bw6 + TCE3 core/non-core | | **Solid organ transplant** | 4-field (2-digit:2-digit) | Eplet-level epitope match (HLAMatchmaker, PIRCHE-II) | | **ICI neoantigen prediction** | 4-field class I + II | NetMHCpan-4.1 minimum | | **HLA-disease association** | 4-field | Standard for GWAS HLA fine-mapping | | **HLA-B\*57:01 abacavir screen** | 4-field, specific | Other \*57 alleles (\*57:03) do NOT cause HSS | | **HLA-B\*15:02 carbamazepine** | 4-field, specific | \*15:02 only; \*15:01 (NFE-common) is not the risk allele | ## G-Groups vs P-Groups: Routinely Confused - **G-groups** collapse alleles with identical DNA sequence across the antigen-recognition exons (class I exons 2-3; class II exon 2). Use for sequence-level lab QC. - **P-groups** collapse alleles encoding identical mature protein across class I positions 1-90 (or class II beta1 domain positions 1-94). Use for epitope-based matching and neoantigen prediction. ## DRB1 + DRB3/4/5 Linkage: The Mandatory Sanity Check DR haplotype linkage is fixed and is the canonical sanity check on any DR typing: | DRB1 allele family | Linked DRB3/4/5 | |---------------------|-----------------| | DR1 (\*01), DR8 (\*08), DR10 (\*10) | None | | DR3 (\*03), DR11 (\*11), DR12 (\*12), DR13 (\*13), DR14 (\*14) | DRB3 | | DR4 (\*04), DR7 (\*07), DR9 (\*09) | DRB4 | | DR15 (\*15), DR16 (\*16) | DRB5 | Any caller reporting DRB4 with `DRB1*15:01` is broken or has a chimera. Use this as a routine QC check on automated pipelines. ## Algorithmic Taxonomy: Short-Read Tools | Tool | Class I | Class II | KIR | Resolution | Approach | Fails when | |------|---------|----------|-----|-----------|----------|-----------| | **OptiType** (Szolek 2014 *Bioinformatics* 30:3310) | Yes (~97% 4-digit) | No | No | 4-field | ILP on exons 2-3 | Class II needed; very deep contamination | | **Polysolver** (Shukla 2015 *Nat Biotechnol* 33:1152) | Yes (~95% 4-digit) | No | No | 4-field | Allele-specific ref alignment | Class II; non-European ancestry under-typing | | **HLA-LA** (Dilthey 2019 *Bioinformatics* 35:4394) | Yes (~94% class I) | Yes (strong class II) | No | 4-field | Graph-based PRG | High RAM/disk (~30-100 GB scratch) | | **T1K** (Song 2023 *Genome Res*) | Yes (~99% 4-digit) | Yes (~99%) | Yes (KIR + KIR3DL2 ligand) | 4-field | EM on consensus reference | Newer; less benchmarking on edge cases | | **HLA-HD** (Kawaguchi 2017 *Hum Mutat* 38:788) | Yes (~98%) | Yes (~95%) | No | 4-field | Bowtie2 against IPD-IMGT | License required for commercial use | | **arcasHLA** (Orenbuch 2020 *Bioinformatics* 36:33) | Yes (~100% 2-field) | Yes (>99% 2-field) | No | 4-field from RNA-seq | EM on STAR alignment | DNA-seq; population prior bias in non-EUR | | **PHLAT, HLAforest, HLAminer, seq2HLA, HLAreporter** | Yes | Some | No | Mostly 2-4 field | Various | Older; superseded | **Operational benchmark consensus:** in the Claeys 2023 *BMC Genomics* 13-tool benchmark (Matey-Hernandez 2018), HLA-HD was the top class-II caller and OptiType (WES) / arcasHLA (RNA) the class-I anchors. T1K (Song 2023, not in that benchmark) adds class I + II + KIR co-typing in one pass and is the 2024-2026 all-rounder recommendation for WGS/WES. ## Long-Read and Ultra-High-Resolution | Tool | Platform | Resolution | Use case | |------|----------|-----------|----------| | **StarPhase** (PacBio official 2024+) | PacBio HiFi | 8-field (full-field) | Transplant-grade typing | | **HLA*ASM** | PacBio HiFi | 8-field | Assembly-based | | **FuFiHLA** (2025 bioRxiv) | PacBio HiFi + ONT R10 | 8-field | Platform-agnostic | | **HLAminer streaming** (Warren 2025) | ONT long-read | 4-field | Streaming nanopore | | **pbaa + StarPhase** | PacBio amplicon | 8-field | Cost-effective targeted typing | ONT R9 was historically unreliable for null-allele discrimination due to homopolymer errors; R10.4 with duplex closes the gap for class I and is competitive with PacBio HiFi for class II. PacBio HiFi remains the gold standard for DPB1 4-field typing. ## SNP-Based HLA Imputation: The Ancestry Footgun When only SNP-array genotypes are available (GWAS cohorts), use imputation: | Tool | Approach | Reference panel | Best for | |------|----------|----------------|----------| | **HIBAG** (Zheng 2014 *Pharmacogenomics J* 14:192) | Random forest from SNP-array | Pre-fit per-ancestry classifiers (EUR, AS, AFR, HIS) | Population-stratified GWAS | | **HLA-TAPAS** (Luo 2021 *Nat Genet* 53:1504) | Multi-ancestry imputation | 21,546 multi-ancestry reference | Cross-ancestry GWAS | | **HLA*IMP:02** (Dilthey 2013) | Hidden Markov | EUR-only | Legacy; EUR-only | | **SNP2HLA** (Jia 2013) | Beagle-based | Type 1 Diabetes / EUR | Older; EUR-only | | **CookHLA** (Cook 2021) | Hybrid SNP2HLA + supplementary | Multi-ancestry refs | Modern alternative to SNP2HLA | | **Multi-Ethnic Reference Panel** (Degenhardt 2019) | Multi-ancestry imputation | Cross-population samples | Cross-ancestry GWAS | **Critical caveat:** imputation panel quality is the limiting factor, NOT the imputation algorithm. EUR-trained HIBAG on East-Asian SNP-array data produces confidently wrong calls. African-ancestry imputation accuracy drops 10-20 percentage points without an ancestry-matched panel (Douillard 2024 *HLA*). For populations underrepresented in IPD-IMGT/HLA itself, imputation is fundamentally limited regardless of method. ## Decision Tree by Scenario | Scenario | Recommended path | Why | |----------|------------------|-----| | WGS/WES, class I only, max speed | OptiType | Best class-I accuracy, ILP-based, fast | | WGS/WES, class I + II, general-purpose | T1K | Best all-rounder; class I + II + KIR co-typing | | WGS/WES, class II reference grade | HLA-LA | Strong class-II accuracy (graph-based PRG) | | RNA-seq tumor/normal for ICI | arcasHLA | RNA-seq native; expressed-allele-aware | | Transplant 6+ field resolution | StarPhase (PacBio HiFi) | 8-field native; reference standard | | Cost-effective targeted typing | pbaa + StarPhase amplicons | Lower cost than WGS | | TCGA-style cancer cohort | Polysolver | TCGA convention; reproduces published values | | SNP array (e.g., UKB) | HIBAG with population-matched panel | No sequencing data | | Multi-ancestry GWAS | HLA-TAPAS | Cross-ancestry reference | | Class II DPB1 4-field certainty | StarPhase or HiFi | Pre-2021 WES kits under-cover DPB1 | | ONT-only data | T1K or HLAminer streaming for class I; ONT R10.4+ duplex for class II | R9 unreliable for nulls | ## HLA and Pharmacogenomics | HLA allele | Drug | Reaction | Population enrichment | OR | |------------|------|----------|----------------------|-----| | **B\*57:01** | Abacavir | Hypersensitivity syndrome | All ancestries (5-8% NFE) | ~100 | | **B\*15:02** | Carbamazepine, oxcarbazepine | SJS/TEN | Han Chinese, Thai, Malay (>=5%) | ~2500 | | **B\*58:01** | Allopurinol | SJS/TEN | Han Chinese, Korean, Thai | ~580 | | **A\*31:01** | Carbamazepine | DRESS, MPE | Europeans, Japanese | ~12 | | **B\*13:01** | Dapsone | DDS | Han Chinese, SE Asian | -- | | **B\*35:02** (NOT \*35:01) | Minocycline | DILI | All ancestries | -- | | **B\*35:01** | TMP-SMX | DILI | Mixed | -- | | **B\*14:01** | TMP-SMX | DILI | African | -- | | **A\*33:01/03** | Terbinafine | DILI | Multi-ancestry | -- | | **DRB1\*15:01 + DQB1\*06:02 haplotype** | Amoxicillin-clavulanate | DILI | Europeans | -- | | **B\*15:13** | Phenytoin | SJS | Malaysian | -- | **Operational rule:** Pharmacogenomic HLA screening requires 4-field resolution; 2-field (e.g., "B*15") misses the specific allele. ## Standard Workflow: T1K on WGS/WES **Goal:** Type HLA class I, class II, KIR from short-read sequencing with KIR3DL1 Bw4/Bw6 ligand prediction. **Approach:** Extract MHC-region reads, run T1K with IPD-IMGT/HLA reference; T1K outputs allele-pair calls + class II haplotype + KIR. ```bash # Extract chr6:28-34 Mb plus alt contigs (alt-aware alignment is critical) samtools view -b -h input.bam chr6:28000000-34000000 chr6_GL000250v2_alt chr6_GL000251v2_alt \ chr6_GL000252v2_alt chr6_GL000253v2_alt chr6_GL000254v2_alt \ chr6_GL000255v2_alt chr6_GL000256v2_alt > hla_region.bam samtools sort -n hla_region.bam -o hla_sorted.bam samtools fastq -1 hla_R1.fq -2 hla_R2.fq -s singletons.fq -0 /dev/null hla_sorted.bam # Run T1K (preset hla; includes class I + II). # Some releases ship the entry point as `run-t1k` (a wrapper script) rather than `t1k`; # verify with `which run-t1k` / `which t1k` before scripting. t1k --preset hla \ -1 hla_R1.fq -2 hla_R2.fq \ -f hla_idx/hlaidx_rna_seq.fa \ -o sample_hla \ --threads 8 # Output: sample_hla_genotype.tsv with HLA-A, B, C, DRB1, DRB3/4/5, DQA1, DQB1, DPA1, DPB1 ``` ## OptiType for Class I (TCGA-Compatible) **Goal:** Type HLA-A, B, C at 4-field from WES with high accuracy. **Approach:** Razers3-based alignment to IMGT class-I reference; ILP optimization to assign reads to allele pairs. ```bash samtools view -h input.bam chr6:28000000-34000000 | samtools fastq -1 R1.fq -2 R2.fq - OptiTypePipeline.py -i R1.fq R2.fq -d -o optitype_out -c config.ini ``` ```ini # config.ini [mapping] razers3=/usr/bin/razers3 threads=8 [ilp] solver=glpk threads=8 [behavior] deletebam=true unpaired_weight=0 use_discordant=false
在 GitHub 查看
这个 SKILL.md 很大,SkillsMP 这里只预览前一段内容。 在 GitHub 查看