- name
- bio-crispr-screens-library-design
- description
- Designs pooled sgRNA libraries for CRISPR knockout, interference (CRISPRi), activation (CRISPRa), Cas12a multiplex, base-editor, and prime-editor screens. Covers on-target scoring (Rule Set 2, Azimuth, DeepSpCas9, CRISPRon), off-target scoring (CFD, MIT), TSS-relative positioning for CRISPRi/a (Horlbeck, Dolcetto, Calabrese), PAM-variant chemistries, control-guide composition, oligo cloning architecture, and library QC. Use when choosing a genome-wide library (GeCKOv2 vs Avana vs Brunello vs TKOv3 vs Inzolia), designing a focused or paralog-focused custom library, picking CRISPRi vs CRISPRa TSS windows, deciding control-guide proportions, or diagnosing library skew and dropout in a freshly cloned pool.
- tool_type
- mixed
- primary_tool
- CRISPOR
## Version Compatibility
Reference examples tested with: CRISPOR 5.01+, BioPython 1.83+, pandas 2.2+, numpy 1.26+, Azimuth 2.0+ (Doench 2016), CRISPRon 1.0+ (Xiang 2021), DeepSpCas9 1.0+ (Kim 2019).
Before using code patterns, verify installed versions match. If versions differ:
- CLI: `crispor.py --help` from the crisporWebsite clone
- Python: Azimuth has no console script; call `azimuth.model_comparison.predict(...)`
If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
## sgRNA Library Design
**"Design a CRISPR library for my screen"** -> Pick a chemistry (Cas9 KO, CRISPRi, CRISPRa, Cas12a, base or prime editor), score candidate guides for on-target activity and off-target liability, position them relative to gene/TSS, add appropriate controls, lay out the oligo for synthesis, and validate the cloned pool.
- Python: `crispor.py` (web + CLI) for batch genome-wide guide scoring with CFD+MIT off-target
- Python: `azimuth` (Microsoft Research) for Rule Set 2 on-target predictions (Brunello-style)
- Python: `CRISPRon`, `DeepSpCas9` for modern deep-learning predictors
- R: `crisprDesign` (Bioconductor) for integrated annotation-aware design
## Library Chemistry Decision Tree
| Goal | Chemistry | Canonical library | Guides/gene | TSS / target window |
|------|-----------|-------------------|-------------|---------------------|
| Loss-of-function essentiality, fitness | SpCas9 KO | Brunello, TKOv3, Avana | 4 (Brunello), 4 (TKOv3), 6 (Avana) | Constitutive exons, prefer aa 5-65% from N-terminus |
| Knockdown of non-cuttable genes, dosage-sensitive | dCas9-KRAB (CRISPRi) | Dolcetto, Horlbeck v2 | 6 (Dolcetto), 5 (Horlbeck) | Optimum +25 to +75 downstream of FANTOM5 TSS, searched out to -50/+300 (Dolcetto, Sanson 2018); -25 to +500 (Horlbeck v2) |
| Gain-of-function, gene activation | dCas9-VP64 / SAM / SunTag (CRISPRa) | Calabrese, Horlbeck-CRISPRa | 6 (Calabrese), 5 (Horlbeck) | -150 to -75 from TSS (Calabrese); -550 to -25 (Horlbeck v2) |
| Paralog buffering, GI screens | enAsCas12a multiplex | Inzolia, in4mer | 4-guide arrays | Constitutive exons |
| Variant function, SNV scanning | CBE / ABE | Custom tiling library | Tile editing windows | Editing window pos 4-8 from PAM-distal end |
| Precise edit, indel-free | Prime editor | Custom PRIDICT-designed | Tile pegRNAs | Anywhere with NGG PAM within 30 nt of edit |
**Fails when:**
- CRISPRi/a targeting wrong TSS: any TSS without FANTOM5 CAGE evidence is suspect; guides positioned against the wrong TSS lose most of their knockdown.
- Cas9 KO of essential paralogs: single-KO buffering hides paralog-redundant essentials (42% of constitutively expressed genes never score, Dede 2020); switch to Cas12a multiplex.
- Base editor over an exon-intron boundary: editing-window bystanders create splice variants instead of the intended SNV.
## On-Target Scoring: Algorithmic Taxonomy
| Predictor | Year | Training set | Strengths | Fails when |
|-----------|------|--------------|-----------|------------|
| Doench Rule Set 1 | 2014 | Flow-sorted GFP+ knockouts | Simple, interpretable | Limited training data; sub-optimal at >NGG context |
| Doench Rule Set 2 / Azimuth 2.0 | 2016 | 1,841 flow-cytometry guides (Doench 2014) plus new guides tiling additional genes | Gold-standard for SpCas9; basis of Brunello | Trained on dropouts; under-predicts efficacy for nuclear-localized targets |
| DeepSpCas9 | 2019 | 12,832 synthetic targets integrated in HEK293T | Spearman ~0.77 vs measured indel frequency on held-out data | Black-box; sensitive to chromatin/context features it wasn't trained on |
| CRISPRon | 2021 | High-throughput indel sequencing | Best for therapeutic-grade target nomination | Slow per-guide; over-fits to its specific cell line |
| DeepHF | 2019 | ~171k guides in HEK293T (WT 55,604; eSpCas9(1.1) 58,167; HF1 56,888) | Separate model per enzyme variant, including WT | Pick the model matching the enzyme actually used |
**Reconciliation:** When predictors disagree, prefer the model whose training cell line matches the screen line (DeepSpCas9 was trained on synthetic targets integrated in HEK293T). For Brunello selection, Azimuth/Rule Set 2 is sufficient because the library was built with it -- introducing a different scorer creates apples-to-oranges ranking with the original library.
## Off-Target Scoring
| Score | Year | Math | Cutoff convention |
|-------|------|------|--------------------|
| MIT (Hsu) | 2013 | Position-weighted mismatch penalty | Specificity score 0-100, higher is better; CRISPOR treats >=50 as a good guide |
| CFD (Doench) | 2016 | Position+nucleotide-specific penalty fit on Brunello | Per-site CFD >0.2 counts a candidate off-target (Doench 2016); CRISPOR's aggregate CFD specificity score is 0-100, higher is better |
| Elevation | 2018 | ML on CFD + mismatch positions | Tighter than CFD |
CFD remains the default for genome-wide library design. **Critical pitfall:** CFD penalizes only mismatches, not bulges; for ≤1 mismatch + 1-bp bulge off-targets, validate empirically with GUIDE-seq or CIRCLE-seq. CRISPOR reports both the MIT (Hsu) and CFD guide specificity scores in a single output.
## Score and Rank sgRNAs for a Target Gene
**Goal:** Generate ranked sgRNA candidates for a single gene, jointly scored on on-target activity (Rule Set 2 / Azimuth) and off-target liability (CFD).
**Approach:** Identify all PAM-adjacent 20-nt protospacers in the target gene's coding sequence, retain only those in the first 5-65% of the protein (constitutive-exon convention from Brunello), filter on GC 30-70% and absence of poly-T (≥4 Ts terminates U6), call Azimuth for on-target and CRISPOR for off-target, and select the top N satisfying both criteria.
```python
import re
import pandas as pd
import numpy as np
from Bio.Seq import Seq
def find_sgrna_candidates(cds_sequence, pam='NGG', guide_length=20):
'''Return all protospacer candidates with PAM coordinates on + strand.
Caller must filter by exon position and Azimuth/CFD score.'''
pam_pattern = re.compile(f'(?=([ACGT]{{{guide_length}}}{pam.replace("N", "[ACGT]")}))')
candidates = []
for strand, seq in [('+', cds_sequence), ('-', str(Seq(cds_sequence).reverse_complement()))]:
for m in pam_pattern.finditer(seq):
spacer = m.group(1)[:guide_length]
if 'TTTT' in spacer or spacer.count('G') + spacer.count('C') not in range(6, 15):
continue
candidates.append({'spacer': spacer, 'strand': strand,
'pos_in_cds': m.start() if strand == '+' else len(seq) - m.start() - 23,
'gc_frac': (spacer.count('G') + spacer.count('C')) / guide_length})
return pd.DataFrame(candidates)
def annotate_exon_position(candidates_df, cds_length):
'''Filter to protospacers within first 5-65% of CDS (Brunello convention).
Reason: N-terminal indels truncate protein; very-N-terminal hits alt initiation;
C-terminal hits miss functional domains (Doench 2016 Nat Biotech).'''
lo, hi = 0.05 * cds_length, 0.65 * cds_length
return candidates_df[(candidates_df['pos_in_cds'] >= lo) & (candidates_df['pos_in_cds'] <= hi)].copy()
```
## CRISPRi / CRISPRa TSS Targeting
**Goal:** Position guides relative to the empirical TSS for maximum knockdown (CRISPRi) or activation (CRISPRa).
**Approach:** Resolve TSS from FANTOM5 CAGE peaks (highest-ranked peak per gene; fall back to Ensembl/RefSeq if absent), define the modality-specific window, score candidate spacers in that window with Rule Set 2 plus the Horlbeck/Sanson CRISPRi/a-tailored rules, and select 5-6 guides per gene biased toward the window center.
```python
def crispri_window(tss_coord, strand='+'):
'''Dolcetto convention: search -50 to +300 around the FANTOM5 highest-rank CAGE peak.
Reason: Sanson 2018 found +25 to +75 nt downstream of the TSS optimal for CRISPRi,
so rank candidates toward that band; the search is relaxed outward to fill the
per-gene guide quota when poorly-annotated TSSs leave too few candidates.'''
if strand == '+':
return (tss_coord - 50, tss_coord + 300)
return (tss_coord - 300, tss_coord + 50)
def crispra_window(tss_coord, strand='+'):
'''Calabrese convention: -150 to -75 upstream of TSS.
Reason: dCas9-VP64 (and SAM, SunTag) activate maximally when bound
just upstream of Pol II loading. Horlbeck v2 CRISPRa uses -550 to -25
(broader, lower per-guide signal). For SAM, prefer Calabrese tightness;
for SunTag, Horlbeck width is acceptable.'''
if strand == '+':
return (tss_coord - 150, tss_coord - 75)
return (tss_coord + 75, tss_coord + 150)
```
**Critical nuance:** Cell-type-specific TSSs differ from the FANTOM5 consensus in ~15% of genes. For tissue-specific screens (e.g., neuron, hepatocyte), re-derive TSSs from a matched CAGE / GRO-seq / PRO-seq dataset before locking guide positions, or knockdown efficiency drops several-fold. The single most common cause of "weak" CRISPRi hits is mis-positioned guides against an alternative TSS.
## Genome-Wide Library Selection
| Library | Year | Modality | Size (genes x guides) | sgRNA rules | Notable |
|---------|------|----------|-----------------------|-------------|---------|
| GeCKOv2 | 2014 | Cas9 KO | ~19k x 6 (~123k) | Exon position + off-target specificity (predates Rule Set 1) | Older; legacy datasets still use it |
| Avana | 2016 | Cas9 KO | 110,257 as published; DepMap screens a ~4-guide subset (Meyers 2017: 70,086 after filtering, 17,670 genes) | Rule Set 1 | Still the Broad's primary Cas9 library; CERES->Chronos changed in 2021, not the library |
| Brunello | 2016 | Cas9 KO | ~19k x 4 (~77k) | Rule Set 2 + CFD | Modern standard for new screens |
| TKOv3 | 2017 | Cas9 KO | ~18k x 4 (~71k) | Hart on/off-target | Bagel/BAGEL2-optimized |
| Humagne | 2020 | enAsCas12a | ~19.8k x 1 dual-guide construct (~20k per set) | enAsCas12a rules | Compact Cas12a sets C and D |
| Horlbeck CRISPRi v2 | 2016 | dCas9-KRAB | ~18k x 5 (~104k) | Horlbeck CRISPRi rules | First-gen, still widely used |
| Dolcetto | 2018 | dCas9-KRAB | ~19k x 3 per set (114,061 across Sets A+B) | Horlbeck + Rule Set 2 | Modern CRISPRi standard |
| Horlbeck CRISPRa | 2016 | dCas9-VP64 | ~18k x 5 (~104k) | Horlbeck CRISPRa rules | Original CRISPRa |
| Calabrese | 2018 | dCas9-VP64 | ~18.9k x 3 per set (113,238 across Sets A+B) | Tight TSS window | Modern CRISPRa standard |
| Inzolia | 2024 | enAsCas12a | ~49k arrays: 19,687 genes (2 arrays each) plus ~4,435 paralog pairs | enAsCas12a rules | Paralog-pair multiplex; ~30% smaller than a typical Cas9 library |
| in4mer | 2024 | Cas12a (4-guide) | Custom | enAsCas12a multiplex | Triple/quadruple KO per cassette |
**dAUC trajectory (essentiality benchmark):** GeCKOv2 < Avana < Brunello/TKOv3 (Doench 2016 + Hart 2017). Moving from 4 to 6 sgRNAs/gene gives diminishing returns; the larger gain is moving from Rule Set 1 to Rule Set 2.
**Cost-coverage tradeoff:** A 77k-guide Brunello at 500x cells/sgRNA needs 38.5M cells in pool, scalable. A 117k-guide Calabrese at 500x needs 59M cells -- often the deciding factor against CRISPRa for difficult-to-grow lines.
## PAM Variants and Alternative Cas Enzymes
| Enzyme | PAM | Spacer length | Best for |
|--------|-----|---------------|----------|
| SpCas9 (WT) | NGG | 20 nt | Standard pooled screens; broadest library support |
| eSpCas9, SpCas9-HF1 | NGG | 20 nt | Lower off-target rate; use for therapeutic-grade nomination |
| SpCas9-NG | NG | 20 nt | Expanded targeting (~4x coverage); accept lower activity per guide |
| SpRY | NRN / NYN | 20 nt | Near-PAMless; coverage at every position; ~50% lower per-guide activity |
| SaCas9 | NNGRRT | 21 nt | AAV-packageable (small ORF); rarely used in pooled screens |
| AsCas12a, LbCas12a | TTTV | 23 nt | AT-rich regions; staggered cut; lower expression noise |
| enAsCas12a (DeWeirdt 2021) | Expanded TTTV + several non-canonical | 23 nt | Combinatorial / paralog screens |
**Decision rule:** If the screen requires every possible TSS position (saturation tiling, dense regulatory dissection), use SpRY despite lower activity; otherwise, NGG is best because the on-target predictors were trained on it.
## Control Guides
A genome-wide library should include:
| Control type | Count | Purpose |
|--------------|-------|---------|
| Non-targeting (scrambled, no genomic match) | 500-1,000 (~1% of library) | Primary null distribution for CRISPRi/a; safe baseline for normalization |
| Safe-harbor (AAVS1, ROSA26-equivalent) | 50-100 | Cas9-only: absorbs cut-toxicity baseline (matters for amplicon-correction) |
| Olfactory receptors (presumed non-expressed) | 50-100 | Second null set for orthogonal normalization |
| Reference essentials (CEGv2 subset: e.g. RPS3, RPL11, EIF3A, POLR2A) | 50-100 | Internal positive control; QC dropout signal |
| Reference non-essentials (NEGv1 subset) | 50-100 | Internal negative control; BAGEL2 calibration |
**Critical pitfall:** Using only AAVS1 as the negative control in a Cas9 screen creates a normalization baseline biased toward "any cut is bad." Always add NTCs or non-essentials so that downstream median normalization and PR-AUC against CEGv2 work without baseline-shift artifacts.
## Library Composition for Specialized Screens
**Paralog buffering (Cas12a multiplex):** Build 4-guide arrays where positions 1-2 target gene A and positions 3-4 target paralog gene B. Inzolia covers ~4,435 paralog pairs within ~49k arrays. Singleton controls (gene A alone, gene B alone) must be included to score genetic interaction = double_KO_LFC - sum(single_KO_LFC).
**Base editor screens (tiling-library design):** Tile NGG-adjacent spacers across exons; ensure editing window (positions 4-8 from PAM-distal end) lands inside coding exons; flag bystander Cs/As in the window for downstream interpretation. Restrict to 50-90% editing efficiency a priori (filter out predicted low-efficacy guides) -- see [[base-editing-analysis]].
**Tiling / regulatory dissection:** Dense (every 5-10 bp) CRISPRi or CRISPRa guides across the candidate region; CRISPRi has broader signal width (good for enhancer discovery) but Cas9-indel tiling has sharper resolution (good for pinpointing critical bases). Pair with CRISPR-SURF deconvolution.
## Oligo Design for Pooled Synthesis
**Goal:** Generate the final oligo sequence ready for chip-based synthesis. Vendor limits differ: Twist oligo pools cap at ~300 nt per oligo with no fixed pool size, GenScript's 92K format spans 20-170 nt, and Agilent OLS 244K spans 30-230 nt.
**Approach:** Add subpool PCR primers (so multiple sublibraries can share a synthesis array), the BsmBI/Esp3I overhang for golden-gate cloning into LentiGuide-Puro (Addgene 52963) or LentiCRISPRv2, and append the tracrRNA scaffold if the array length permits.
```python
def build_oligo(spacer, vector='lentiGuide-Puro', subpool_idx=None):
'''Construct final oligo for pooled synthesis.
LentiGuide-Puro / LentiCRISPRv2 use BsmBI (Esp3I) with these overhangs:
forward: 5'-CACCG[spacer]-3'
reverse: 5'-AAAC[revcomp(spacer)]C-3'
For chip synthesis, the spacer is flanked by subpool-specific PCR primers.'''
subpool_fwd = {
1: 'GGAAAGGACGAAACACCG', # subpool 1 forward primer + BsmBI overhang
2: 'GAGGCACTGGGCAGGTACCG',
}.get(subpool_idx, 'GGAAAGGACGAAACACCG')
# First 33 nt of the Chen 2013 sgRNA(F+E) optimized scaffold. NOTE: lentiGuide-Puro (#52963)
# and lentiCRISPRv2 (#52961) carry the ORIGINAL scaffold; F+E belongs to lentiCRISPRv2-Opti (#163126).
View on GitHub