| name | rfdiffusion |
| description | > Use when this capability is needed. |
RFdiffusion Backbone Generation
Prerequisites
| Requirement | Minimum | Recommended |
|---|
| Python | 3.9+ | 3.10 |
| CUDA | 11.7+ | 12.0+ |
| GPU VRAM | 16GB | 24GB (A10G) |
| RAM | 16GB | 32GB |
How to run
First time? See Installation Guide to set up Modal and biomodals.
Option 1: Modal (recommended)
git clone https://github.com/hgbrian/biomodals && cd biomodals
modal run modal_rfdiffusion.py \
--pdb target.pdb \
--contigs "A1-150/0 70-100" \
--hotspot "A45,A67,A89" \
--num-designs 100
GPU=A100 TIMEOUT=60 modal run modal_rfdiffusion.py \
--pdb target.pdb \
--contigs "A1-150/0 70-100" \
--num-designs 100
GPU: A10G (24GB) | Timeout: 30min default
Option 2: Local installation
git clone https://github.com/RosettaCommons/RFdiffusion.git
cd RFdiffusion && pip install -e .
wget http://files.ipd.uw.edu/pub/RFdiffusion/models/Complex_base_ckpt.pt
python run_inference.py \
inference.input_pdb=target.pdb \
contigmap.contigs=[A1-150/0 70-100] \
ppi.hotspot_res=[A45,A67,A89] \
inference.num_designs=100
Config Schema (Hydra)
Contigmap Syntax
contigmap.contigs=[50-100]
contigmap.contigs=[A1-150/0 70-100]
contigmap.contigs=[20-40/0 A10-30/0 20-40]
contigmap.contigs=[A1-100/0 B1-100/0 60-80]
contigmap.contigs=[A1-150/0 50-100]
Hotspot Specification
ppi.hotspot_res=[A45,A67,A89]
Common mistakes
Contig Syntax
✅ Correct:
contigmap.contigs=[A1-150/0 70-100]
❌ Wrong:
contigmap.contigs=[A1-150 70-100]
contigmap.contigs="A1-150/0 70-100"
contigmap.contigs=[A1-150/0, 70-100]
Hotspot Residues
✅ Correct:
ppi.hotspot_res=[A45,A67,A89]
❌ Wrong:
ppi.hotspot_res=[45,67,89]
ppi.hotspot_res=[A45, A67, A89]
ppi.hotspot_res="A45,A67,A89"
Complete Parameter Reference
Core Parameters
| Parameter | Default | Range | Description |
|---|
inference.num_designs | 10 | 1-10000 | Number of designs to generate |
inference.input_pdb | - | path | Target structure file |
inference.output_prefix | output | string | Output filename prefix |
diffuser.T | 50 | 20-200 | Diffusion timesteps |
denoiser.noise_scale_ca | 1.0 | 0.0-2.0 | CA atom noise (0.5-0.8 = conservative) |
denoiser.noise_scale_frame | 1.0 | 0.0-2.0 | Frame noise |
inference.ckpt_override_path | - | path | Model checkpoint |
potentials.guide_scale | 1.0 | 0.1-10 | Guidance strength |
potentials.guide_decay | constant | string | Decay type |
Advanced Parameters
| Parameter | Default | Description |
|---|
diffuser.partial_T | None | Start diffusion from timestep T (partial diffusion) |
contigmap.inpaint_str | None | Sequence positions to inpaint |
scaffoldguided.scaffoldguided | false | Enable scaffold-guided generation |
scaffoldguided.target_pdb | None | Scaffold template PDB |
ppi.binderlen | None | Specify exact binder length |
Symmetry Parameters
| Parameter | Default | Description |
|---|
symmetry.symmetry | None | Symmetry type (C2, C3, C4, D2, etc.) |
symmetry.recenter | true | Recenter symmetric assembly |
symmetry.radius | None | Radius constraint for symmetric assembly |
Fold Conditioning
| Parameter | Default | Description |
|---|
contigmap.provide_seq | None | Provide sequence for fold conditioning |
contigmap.inpaint_seq | None | Positions for sequence inpainting |
Model Checkpoints
| Checkpoint | Use Case |
|---|
Complex_base_ckpt.pt | Binder design (default) |
Base_ckpt.pt | De novo monomers |
ActiveSite_ckpt.pt | Active site scaffolding |
InpaintSeq_ckpt.pt | Sequence inpainting |
Common workflows
Binder Design
- Prepare target PDB (trim to binding region + 10A buffer)
- Identify 3-6 hotspot residues (exposed, conserved)
- Generate 100-500 backbones
- Pass to proteinmpnn for sequence design
Motif Scaffolding
- Extract motif coordinates
- Use
/0 to fix motif in contigmap
- Generate surrounding scaffold
- Validate motif preservation (RMSD < 1.5A)
Symmetric Oligomers
python run_inference.py \
symmetry.symmetry=C3 \
contigmap.contigs=[100-150] \
inference.num_designs=50
python run_inference.py \
symmetry.symmetry=D2 \
contigmap.contigs=[80-120] \
symmetry.radius=25
Partial Diffusion (Refinement)
python run_inference.py \
inference.input_pdb=initial.pdb \
diffuser.partial_T=10 \
contigmap.contigs=[A1-100]
Output format
output/
├── output_0.pdb # Generated backbone
├── output_1.pdb
├── ...
└── output_99.pdb
Each PDB contains polyalanine backbone - use proteinmpnn for sequence.
Sample output
Successful run
$ python run_inference.py inference.input_pdb=target.pdb contigmap.contigs=[A1-150/0 70-100] inference.num_designs=100
[INFO] Loading model from Complex_base_ckpt.pt
[INFO] Generating design 1/100...
[INFO] Generating design 50/100...
[INFO] Generating design 100/100...
[INFO] Saved 100 designs to output/
Generated:
output/output_0.pdb (85 residues)
output/output_1.pdb (92 residues)
...
What good output looks like:
- File size: 3-8 KB per PDB (backbone only)
- Residue count within specified range
- Secondary structure visible in PyMOL (helices/sheets, not random coil)
Decision tree
Should I use RFdiffusion?
│
├─ Need to generate protein backbone?
│ ├─ Yes → Continue below
│ └─ No, already have backbone → Use ProteinMPNN
│
├─ What type of design?
│ ├─ Binder for protein target → RFdiffusion ✓
│ ├─ De novo monomer → RFdiffusion ✓
│ ├─ Motif scaffolding → RFdiffusion ✓
│ └─ Symmetric assembly → RFdiffusion ✓
│
└─ Priority?
├─ Need highest success rate → Consider BindCraft
├─ Need diversity/exploration → RFdiffusion ✓
└─ Need all-atom precision → Consider BoltzGen
Typical performance
| Campaign Size | Time (A10G) | Cost (Modal) | Notes |
|---|
| 100 backbones | 20-30 min | ~$3 | Quick exploration |
| 500 backbones | 1.5-2h | ~$12 | Standard campaign |
| 1000 backbones | 3-4h | ~$25 | Large campaign |
Expected downstream yield: ~10-15% of backbones pass full QC after sequence design + validation.
Verify
ls output/*.pdb | wc -l
Troubleshooting
Designs lack secondary structure: Decrease noise_scale to 0.5-0.8
Binder not contacting hotspots: Verify residue numbering, increase num_designs
OOM errors: Reduce batch size or use A100 GPU
Slow generation: Reduce diffuser.T to 25-35
Error interpretation
| Error | Cause | Fix |
|---|
RuntimeError: CUDA out of memory | GPU VRAM exceeded | Use A100 or reduce designs per batch |
KeyError: 'A' | Chain not found in PDB | Check chain IDs with grep ^ATOM target.pdb | cut -c22 | sort -u |
ValueError: invalid contig | Syntax error in contigs | Check for spaces, quotes, commas (see Common Mistakes) |
FileNotFoundError: ckpt | Missing model weights | Download from IPD website |
Next: proteinmpnn for sequence design → structure prediction for validation → protein-qc for filtering.
Converted and distributed by TomeVault — claim your Tome and manage your conversions.