- name
- bio-clinical-databases-tumor-mutational-burden
- description
- Calculates tumor mutational burden from WES/WGS/panel data with Friends of Cancer Research harmonization equations, per-assay calibration (FDA 10/Mb = 7.8 TSO500 = 8.4 OncomineTML), synonymous/indel/germline filtering, hypermutator tiering, blood TMB, and integration with HLA-LOH and neoantigen quality (Luksza 2017 fitness). Use when assessing ICI eligibility under tumor-specific cutoffs (McGrail 2021), comparing tissue vs bTMB, or auditing TMB-H reporting against ESMO 2024 and FDA pembrolizumab pan-tumor 2020.
- tool_type
- python
- primary_tool
- cyvcf2
## Version Compatibility
Reference examples tested with: cyvcf2 0.30+, VEP 111+ (or snpEff 5.2+), pandas 2.2+, numpy 1.26+, LOHHLA 1.0+ (McGranahan 2017), DASH 1.0+ (Pyke 2022). v4.1 (May 2024) gnomAD is current for germline subtraction. Friends of Cancer Research TMB harmonization framework (Vega 2021 *Ann Oncol*) and ESMO 2024 (Mosele *Ann Oncol*) define the operational thresholds.
Before using code patterns, verify installed versions match. If versions differ:
- Python: `pip show <package>` then `help(module.function)` to check signatures
- CLI: `<tool> --version`
If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying. TMB calculation requires VCF with VEP / snpEff / Funcotator consequence annotations; the panel size used as denominator MUST match the assay's actual scored region, NOT the panel's total content.
# Tumor Mutational Burden; Calculation, Harmonization, ICI Eligibility
**'Calculate TMB from this somatic VCF and apply ICI eligibility cutoff'** -> Count nonsynonymous coding variants passing VAF/depth/germline filters; divide by assay scored region in Mb; apply assay-calibrated TMB-H cutoff; integrate with MSI / HLA-LOH / neoantigen quality.
- Python: `cyvcf2.VCF()` + VEP/snpEff consequence parsing + panel-size normalization
- CLI: `bcftools view` filtering + custom counting
- HLA-LOH: LOHHLA (McGranahan 2017 *Cell*) or DASH (Pyke 2022 *Nat Commun*)
- Neoantigen quality: pVAC-tools, NetMHCpan-4.1, Luksza 2017 fitness model
## Regulatory and Trial Landscape
| Event | Year | Threshold | Notes |
|-------|------|-----------|-------|
| **KEYNOTE-158 + FDA pembrolizumab pan-tumor approval** | 2020 | TMB-H >= 10 mut/Mb | FoundationOne CDx companion diagnostic; 10 cohorts |
| **Friends of Cancer Research TMB harmonization Phase I (Merino 2020)** | 2020 | -- | 11 panels vs WES truth; 3-fold panel-specific differences |
| **Friends of Cancer Research Phase II (Vega 2021)** | 2021 | Calibration equations | 19 platforms; per-assay calibration to WES-aligned TMB-Mb |
| **ESMO 2024 (Mosele *Ann Oncol*)** | 2024 | TMB-H >= 10/Mb retained (tumour-agnostic, ESCAT IB) | Tumour-type limits per McGrail 2021 |
| **KEYNOTE-189 (NSCLC + pembrolizumab + chemo)** | 2018 | -- | TMB-H did NOT enrich for benefit with chemo backbone |
| **POSEIDON / KEYNOTE-021 / KEYNOTE-407** | 2019-2022 | -- | TMB inconsistent with chemo backbones |
| **B-F1RST + BFAST Cohort C (bTMB)** | 2022 | bTMB >= 16/Mb | BFAST Cohort C FAILED primary endpoint |
## Friends of Cancer Research Harmonization: Cross-Panel Calibration
Merino 2020 *J Immunother Cancer*: in silico panel sampling from TCGA WES truth showed panel-specific TMB can differ 3-fold for identical samples. Vega 2021 *Ann Oncol* derived per-panel calibration equations to translate panel TMB to WES-aligned TMB-Mb.
**Per-panel calibration to FoundationOne 10/Mb sensitivity:**
| Panel | Scored region (Mb) | Equivalent threshold for FDA 10/Mb pan-tumor | Fails when |
|-------|---------------------|----------------------------------------------|-----------|
| **FoundationOne CDx** | 0.8 Mb scored (NOT 1.1 Mb total) | 10 mut/Mb (FDA reference; F1CDx companion) | Using 1.1 Mb panel total inflates TMB ~37%; pipeline excludes synonymous (F1CDx includes them) |
| **MSK-IMPACT v3** | 0.98 Mb | ~10 (full Vega 2021 calibration recommended) | Tumor purity < 30%; non-paired-normal mode |
| **MSK-IMPACT v4** | 1.22 Mb | ~10 | -- |
| **TruSight Oncology 500** | ~1.3 Mb scored (from 1.94 Mb total) | **7.8 mut/Mb** | Pipeline uses 10/Mb instead of the TMB2-calibrated 7.8 (Ramos-Paradas 2021) |
| **Oncomine Tumor Mutation Load** | 1.2 Mb | **8.4 mut/Mb** | Pipeline uses 10/Mb instead of the TMB2-calibrated 8.4 (Ramos-Paradas 2021) |
| **Caris MI Tumor Seek** | ~1.2 Mb |; (verify Caris docs) | -- |
| **Tempus xT v3** | 0.6 Mb | -- | Below 0.8 Mb minimum reliability threshold |
| **Predicine ATLAS** | ~0.6 Mb | -- | Below 0.8 Mb minimum; high sampling variance |
**TMB =/= TMB across vendors.** Manuscripts that compare TMB across panels without per-assay calibration are unreviewable. Use the Vega 2021 calibration equations or WES re-projection.
## Variant-Counting Subtleties
These choices alter TMB by 5-20%:
| Variable | Convention | Notes |
|----------|-----------|-------|
| **Synonymous variants** | **FoundationOne CDx INCLUDES synonymous** (rationale: reduces sampling noise); MSK-IMPACT and most academic pipelines exclude | The FDA companion diagnostic counts synonymous; frequent misconception |
| **Indels** | FoundationOne includes; some assays exclude frameshift only | 5-15% TMB impact |
| **Germline subtraction** | Paired-normal (gold standard); else gnomAD AF <=0.5% (sometimes 1%) for tumor-only | Population-stratified gnomAD AF for ancestry-diverse cohorts |
| **VAF threshold** | FoundationOne >=5%; >=10% for tumor-only no UMI; down to 2% with paired-normal | Lower VAF risks contamination/artifacts |
| **Hotspots** | COSMIC-confirmed driver hotspots typically EXCLUDED (not random) | Inflates TMB if included |
| **Tumor purity** | FoundationOne >=20%; MSK-IMPACT >=30% | Below floor erodes VAF-based filtering |
| **VEP version** | Pin to assay's annotation version | gnomAD v4 uses VEP 105 |
## Hypermutator Tiering
| Class | Threshold | Common etiology |
|-------|-----------|----------------|
| **TMB-H (FDA pan-cancer)** | >= 10 mut/Mb | Variable; ICI eligible |
| **Hypermutator (research)** | >= 100 mut/Mb | MMR-D, POLE-exo |
| **Ultra-hypermutator** | >= 500 mut/Mb | POLE+MMR concurrent |
MSI-H typically 30-50 mut/Mb; pure POLE-exo P286R 100-300 mut/Mb; POLE-exo + MMR-D exceeds 500. MSI-H and TMB-H overlap substantially (~83% of MSI-H are TMB-H) but only ~16% of TMB-H solid tumors are MSI-H (Chalmers 2017 *Genome Med* 9:34).
## The Tumor-Type-Specific Cutoff Debate
**McGrail 2021** *Ann Oncol* is the most damning paper for the universal 10/Mb cutoff. TMB-H predicts ICI response in melanoma, NSCLC, bladder; but FAILS in breast, prostate, glioma. ORR in TMB-H melanoma/NSCLC/bladder was 39.8%; TMB-H breast/prostate/glioma was 15.3%. Mechanistic explanation: TMB only predicts when baseline CD8 T-cell infiltrate is present.
**Sha 2020** *Cancer Discov*: TMB-H predicts ICI benefit in MSS subset but adds nothing on top of MSI-H (because MSI-H is uniformly hypermutator and uniformly responsive).
**Samstein 2019** *Nat Genet* (MSK-IMPACT 1,662 ICI-treated): cancer-specific TMB cutoffs (top 20% within each tumor type) outperform universal 10/Mb.
**ESMO 2024** retained TMB-H >= 10/Mb pan-tumor (tumour-agnostic, ESCAT IB). The tumour-type limits (poor performance in breast, prostate, glioma) come from **McGrail 2021**, not ESMO.
## Blood TMB (bTMB): The Negative-Trial Story
**Gandara 2018** *Nat Med*: bTMB on Foundation Medicine FoundationACT panel; POPLAR + OAK retrospective. bTMB >= 16 mut/Mb showed PFS benefit with atezolizumab in NSCLC.
**B-F1RST (Kim 2022)**: prospective phase 2 test of bTMB >= 16 as a first-line atezolizumab predictor in NSCLC; did NOT meet its pre-specified primary endpoint (bTMB-H improved ORR 28.6% vs 4.4%, only a non-significant PFS/OS trend).
**BFAST Cohort C (Peters 2022)**: FAILED primary endpoint; atezolizumab vs chemo in bTMB-H NSCLC did not improve investigator-assessed PFS. Dominant confounder: low ctDNA shed fraction produces false-negative bTMB.
**Operational state:** bTMB is research-grade in tissue-naive settings; tissue TMB remains the regulatory standard.
## Neoantigen Quality: Beyond Raw TMB
**Luksza 2017** *Nature*: neoantigen fitness model. Combines "non-selfness" (TCR recognition probability via IEDB similarity) + "selfness" (MHC binding affinity differential vs WT peptide). Pancreatic-cancer validation (Balachandran 2017 *Nature*): long-term survivors had higher-quality neoantigens. Luksza 2022 *Nature*: immunoediting over 10 years.
**McGranahan 2016** *Science*: **clonal neoantigen burden** (mutations present in all tumor cells) predicts ICI response better than total. Subclonal-rich tumors evade despite high TMB.
**HLA-LOH** (McGranahan 2017 *Cell*, LOHHLA; Pyke 2022 *Nat Commun*, DASH; Montesion 2021 *Cancer Discov* for the ~17% pan-cancer estimate): HLA-LOH occurs in ~40% of NSCLC and abolishes neoantigen presentation for the lost allele. ~17% pan-cancer; >30% in HNSCC / NSCLC / cervical. Co-occurs with high subclonal burden + APOBEC + immune escape.
## Decision Tree by Scenario
| Scenario | Recommended path | Why |
|----------|------------------|-----|
| Pan-tumor ICI eligibility (FDA pembrolizumab) | TMB-H >= 10/Mb on FoundationOne CDx | FDA companion diagnostic |
| Non-FoundationOne panel | Apply per-assay calibration to the FoundationOne 10/Mb equivalent | TSO500 = 7.8; Oncomine = 8.4 (Ramos-Paradas 2021) |
| WES TMB | Compute directly; threshold per ESMO 2024 = 10/Mb | WES is reference standard |
| Tissue-naive bTMB | Caution: BFAST Cohort C failed | Research-grade; check ctDNA shed fraction |
| Breast / prostate / glioma | TMB-H does not enrich ICI response per McGrail 2021 | Tumor-type-specific cutoffs |
| MSI-H + TMB-H concurrence | MSI-H supersedes for ICI biomarker decision | Sha 2020 |
| Hypermutator characterization (>=100/Mb) | Confirm MMR-D or POLE-exo via signatures + IHC | Co-occurrence is common |
| Neoantigen quality (research) | Luksza fitness + HLA-LOH (LOHHLA / DASH) + clonality (McGranahan) | Beyond raw TMB |
| Cross-panel comparison | Vega 2021 calibration equations OR WES re-projection | Direct comparison invalid |
## Standard Workflow
**Goal:** Compute TMB from a VEP-annotated somatic VCF with full filtering.
**Approach:** Parse cyvcf2; apply VAF + depth + germline (gnomAD) filters; count nonsynonymous coding consequences; divide by scored Mb.
```python
from cyvcf2 import VCF
import re
NONSYNONYMOUS_CONSEQUENCES = {
'missense_variant', 'stop_gained', 'stop_lost', 'start_lost', 'start_retained',
'frameshift_variant', 'inframe_insertion', 'inframe_deletion',
'splice_donor_variant', 'splice_acceptor_variant',
'protein_altering_variant', 'initiator_codon_variant'
}
# Vega 2021-calibrated scored regions (Mb)
PANEL_SCORED_REGION = {
'FoundationOne_CDx': 0.8, # Scored region; NOT 1.1 panel total
'MSK_IMPACT_v3': 0.98,
'MSK_IMPACT_v4': 1.22,
'TSO500': 1.3, # Scored from 1.94 total
'Oncomine_TML': 1.2,
'Caris_MI': 1.2,
'Tempus_xT_v3': 0.6, # Borderline reliability
'WES': 30.0,
'WGS': 3000.0
}
# TMB2 (Ramos-Paradas 2021) equivalent thresholds for FDA 10/Mb FoundationOne sensitivity
ASSAY_TMB_H_CUTOFF = {
'FoundationOne_CDx': 10.0,
'TSO500': 7.8,
'Oncomine_TML': 8.4,
'MSK_IMPACT_v3': 10.0, # Approximate; full Vega 2021 calibration recommended
'MSK_IMPACT_v4': 10.0,
'WES': 10.0
}
def parse_consequences_from_vep(csq_field, csq_header):
'''Parse VEP CSQ INFO field; returns list of per-transcript consequence types.'''
if not csq_field:
return []
cons_idx = csq_header.index('Consequence')
out = []
for transcript in csq_field.split(','):
fields = transcript.split('|')
if len(fields) > cons_idx:
out.append(fields[cons_idx])
return out
def is_nonsynonymous(consequences, include_synonymous=False):
'''Check if variant has nonsynonymous coding consequence.
FoundationOne CDx convention INCLUDES synonymous (set include_synonymous=True).
MSK-IMPACT and most academic pipelines exclude.
'''
target = set(NONSYNONYMOUS_CONSEQUENCES)
if include_synonymous:
target.add('synonymous_variant')
for cons_str in consequences:
for cons in cons_str.split('&'):
if cons in target:
return True
return False
def calculate_tmb(vcf_path, scored_region_mb, csq_header,
min_vaf=0.05, min_depth=100, max_gnomad_af=0.005,
include_synonymous=False, exclude_hotspots=True,
hotspot_bed=None):
'''Calculate TMB with filtering per Vega 2021 harmonization.
Args:
scored_region_mb: panel's SCORED region (NOT total panel)
min_vaf: 0.05 (FoundationOne) to 0.10 (tumor-only no UMI)
max_gnomad_af: 0.005 (0.5%) typical for tumor-only germline filter
include_synonymous: True for FoundationOne CDx-compatible; False for MSK-IMPACT
exclude_hotspots: COSMIC drivers excluded (not random mutations)
'''
vcf = VCF(vcf_path)
cons_idx = csq_header.index('Consequence') if 'Consequence' in csq_header else 1
nonsyn_count = 0
total_pass = 0
for v in vcf:
if v.FILTER is not None: # FILTER == None means PASS in cyvcf2
continue
depth = v.INFO.get('DP', 0)
if depth < min_depth:
continue
vaf = _get_vaf(v)
if vaf is None or vaf < min_vaf:
continue
gnomad_af = v.INFO.get('gnomAD_AF', 0) or v.INFO.get('AF_popmax', 0) or 0
if gnomad_af > max_gnomad_af:
continue
total_pass += 1
csq = v.INFO.get('CSQ', '')
consequences = parse_consequences_from_vep(csq, csq_header)
if is_nonsynonymous(consequences, include_synonymous=include_synonymous):
nonsyn_count += 1
tmb = nonsyn_count / scored_region_mb
return {
'tmb': round(tmb, 2),
'nonsynonymous_count': nonsyn_count,
'total_passing_filters': total_pass,
'scored_region_mb': scored_region_mb
}
def _get_vaf(variant):
'''Extract VAF from genotype FORMAT fields (Mutect2 AD or AF).'''
try:
ad = variant.format('AD')
if ad is not None and len(ad) > 0:
ad0 = ad[0]
total = sum(ad0)
在 GitHub 查看