Performs molecular similarity searching using Tanimoto, Tversky, Dice, and cosine coefficients on bit/count fingerprints with explicit choice rules for symmetric vs asymmetric measures, scaffold-hopping vs lead-optimization regimes, activity-cliff diagnosis, and large-library nearest-neighbor methods (BulkTanimoto, MHFP6 LSH forest, USRCAT). Use when ranking compounds by structural resemblance to a query, clustering libraries, finding analogs, or diagnosing activity cliffs.
Performs molecular similarity searching using Tanimoto, Tversky, Dice, and cosine coefficients on bit/count fingerprints with explicit choice rules for symmetric vs asymmetric measures, scaffold-hopping vs lead-optimization regimes, activity-cliff diagnosis, and large-library nearest-neighbor methods (BulkTanimoto, MHFP6 LSH forest, USRCAT). Use when ranking compounds by structural resemblance to a query, clustering libraries, finding analogs, or diagnosing activity cliffs.
Before using code patterns, verify installed versions match. If versions differ:
Python: pip show <package> then help(module.function) to check signatures
If code throws ImportError, AttributeError, or TypeError, introspect the installed
package and adapt the example to match the actual API rather than retrying.
Similarity Searching
Find structurally similar compounds and cluster libraries by similarity. The choice of similarity coefficient and fingerprint is task-aware: Tanimoto for symmetric similarity in lead optimization, Tversky for asymmetric "substructure-like" queries, Dice for higher sensitivity in low-similarity regimes, and MaxCommon Substructure (MCS) for scaffold-hopping. Tanimoto similarity above 0.7 is not a guarantee of activity preservation; activity cliffs (similar molecules with dissimilar activities) are common (Maggiora 2014).
For fingerprint choice, see chemoinformatics/molecular-descriptors. For 3D shape similarity, see chemoinformatics/shape-similarity.
Similarity Coefficient Taxonomy
Coefficient
Formula
Range
Symmetric
Use case
Fails when
Tanimoto
c / (a + b - c)
0-1
Yes
Default for ECFP4 similarity, ranking analogs
Saturates at low similarity (drug vs natural product)
Dice
2c / (a + b)
0-1
Yes
Bit or nonnegative sparse-count vectors when Dice semantics are intended
Thresholds depend on vector type; analog choice subjective
Mostly noise; use 3D shape or pharmacophore instead
ECFP4 not informative
These bands are working defaults for ECFP4-like fingerprints, not transferable calibration. Inspect the target dataset's similarity distribution and known series before setting a cutoff. Maggiora's similarity principle states "similar molecules tend to have similar activity" -- but activity cliffs (Stumpfe & Bajorath 2012) violate this. Treat high ECFP4 similarity as a prioritization signal, not evidence that activity will be preserved.
Decision Tree by Scenario
Goal
Workflow
Tools
Find analogs of a hit (lead opt)
ECFP4 Tanimoto >=0.7 search
RDKit BulkTanimotoSimilarity
Find scaffold hops
FCFP4 OR AtomPair Tanimoto >=0.5 + filter MCS
RDKit + rdFMCS
Cluster library by chemotype
Butina clustering at Tanimoto 0.6 cutoff
RDKit Butina.ClusterData
Diversity sampling
MaxMin selection on Tanimoto
RDKit rdSimDivPickers.MaxMinPicker
Nearest neighbors in >1M library
LSH forest with MHFP6
mhfp.lsh_forest.LSHForestHelper
Activity cliff diagnosis
Tanimoto + pIC50 delta scatter
Custom analysis
3D similarity (shape)
USRCAT / Open3DAlign / ROCS
shape-similarity skill
Tanimoto Similarity (single query, large library)
Goal: Rank a library by ECFP4 Tanimoto similarity to a query molecule, returning hits above a threshold.
Approach: Generate ECFP4 fingerprints for all molecules once, then use BulkTanimotoSimilarity for O(N) lookup.
from rdkit import Chem, DataStructs
from rdkit.Chem import rdFingerprintGenerator
defprecompute_fps(smiles_list, radius=2, nBits=2048):
generator = rdFingerprintGenerator.GetMorganGenerator(
radius=radius, fpSize=nBits)
fps = []
for smi in smiles_list:
mol = Chem.MolFromSmiles(smi)
if mol isNone:
fps.append(None)
else:
fps.append(generator.GetFingerprint(mol))
return fps
defsearch(query_smi, library_fps, threshold=0.7):
qmol = Chem.MolFromSmiles(query_smi)
if qmol isNone:
raise ValueError('invalid query SMILES')
generator = rdFingerprintGenerator.GetMorganGenerator(radius=2, fpSize=2048)
qfp = generator.GetFingerprint(qmol)
valid = [(i, fp) for i, fp inenumerate(library_fps) if fp isnotNone]
sims = DataStructs.BulkTanimotoSimilarity(qfp, [fp for _, fp in valid])
return [(source_i, sim) for (source_i, _), sim inzip(valid, sims)
if sim >= threshold]
Tversky for Asymmetric Substructure-Like Search
Goal: Rank a library by how much each compound "contains" the features of the query (asymmetric).
Approach: Tversky with alpha=1, beta=0 rewards compounds containing query bits (substructure-like) while ignoring extra bits in the compound.
from rdkit import DataStructs
deftversky_substructure_like(qfp, lib_fps, alpha=1.0, beta=0.0):
return [DataStructs.TverskySimilarity(qfp, f, alpha, beta) for f in lib_fps if f]
Use case: identifying analogs that extend a pharmacophore vs. exact-similarity ranking.
Butina Clustering
Goal: Group a library around Taylor-Butina centroids whose assigned neighbors are within the selected distance cutoff.
from rdkit.ML.Cluster import Butina
defcluster(mols, cutoff=0.4):
generator = rdFingerprintGenerator.GetMorganGenerator(radius=2, fpSize=2048)
fps = [generator.GetFingerprint(m) for m in mols]
n = len(fps)
dists = []
for i inrange(1, n):
sims = DataStructs.BulkTanimotoSimilarity(fps[i], fps[:i])
dists.extend([1 - s for s in sims])
return Butina.ClusterData(dists, n, cutoff, isDistData=True)
cutoff=0.4 means each assigned member was a neighbor of its selected centroid at Tanimoto >= 0.6. It does not guarantee that every pair of non-centroid members has Tanimoto >= 0.6. The first molecule in each returned cluster is the cluster centroid.
Trade-off: Butina materializes O(N^2) pairwise distances. Benchmark memory and runtime on the actual library; for much larger collections, use an approximate method such as an MHFP6 LSH forest.
Diversity Selection (MaxMin)
Goal: Select N diverse compounds from a library by maximizing the minimum pairwise distance.
Use cases: identify scaffold across a series, build scaffold hopping queries, generate consensus pharmacophore.
Limit: MCS search can become combinatorial as input count, size, and structural divergence grow. Set a finite timeout, inspect result.canceled, and consider pre-clustering or reducing the comparison set; no molecule-count or atom-count boundary guarantees tractability.
Activity Cliff Diagnosis
Goal: Detect pairs of similar molecules with dissimilar activities (cliffs).
Approach: Compute pairwise ECFP4 Tanimoto + pIC50 delta. Flag pairs with high similarity and large activity gap.
defactivity_cliffs(df, sim_threshold=0.85, activity_gap=2.0, activity_col='pIC50'):
generator = rdFingerprintGenerator.GetMorganGenerator(radius=2, fpSize=2048)
mols = [Chem.MolFromSmiles(s) for s in df['smiles']]
ifany(mol isNonefor mol in mols):
raise ValueError('activity-cliff input contains invalid SMILES')
fps = [generator.GetFingerprint(mol) for mol in mols]
cliffs = []
for i inrange(len(fps)):
sims = DataStructs.BulkTanimotoSimilarity(fps[i], fps[i+1:])
for j_off, sim inenumerate(sims):
j = i + 1 + j_off
if sim >= sim_threshold:
gap = abs(df[activity_col].iloc[i] - df[activity_col].iloc[j])
if gap >= activity_gap:
cliffs.append((i, j, sim, gap))
return cliffs
Activity cliffs flag (a) measurement noise, (b) cryptic SAR (e.g. ring-flip changing dihedral), (c) protein conformational selection, or (d) actually informative SAR. Cliffs are an opportunity for medchem investigation, not necessarily an error.
For large libraries, direct all-pairs comparison becomes expensive. The mhfp package provides an LSH-forest helper for approximate nearest-neighbor retrieval over MHFP6 fingerprints. Measure recall and latency against an exact subset for the project dataset.
The returned neighbors are approximate in MHFP6 space. Benchmark them against an exact MHFP-distance search on a representative subset before choosing LSH parameters.
Mechanism: A local circular fingerprint may not preserve the distinctions needed for a particular mixed-modality retrieval task.
Symptom: Known relevant neighbors are not enriched above background, or retrieval metrics and neighborhood stability are poor on target-relevant controls. A low mean pairwise similarity alone is not a universal saturation test.
Fix: Benchmark ECFP4 against alternatives such as MHFP6 or MAP4 using held-out analog recovery, scaffold-aware retrieval, or another task-aligned metric. Do not transfer raw-score thresholds between fingerprint families.
Fix: Use approximate clustering (HDBSCAN on UMAP-reduced fingerprints) or LSH-based clustering on MHFP6.
MCS -- exponential timeout
Trigger: Mol set with low overlap, large molecules, or many input mols.
Mechanism: MCS search is NP-hard; algorithm tries every atom-mapping permutation within timeout.
Symptom: Returns small partial MCS or empty result.
Fix: Raise timeout; reduce input mol count; pre-cluster by Tanimoto first then MCS within clusters.
Tanimoto = 1.0 != same molecule
Trigger: Comparing fingerprints between two molecules that hash to the same bits.
Mechanism: A folded hashed fingerprint can map distinct atom environments to the same bits; collision frequency depends on molecule size, radius, and fingerprint length.
Symptom: Two structurally different molecules report Tanimoto 1.0.
Fix: For exact identity, compare canonical SMILES or InChIKey, not fingerprint. Use unhashed sparse fingerprint to disambiguate.
Similarity threshold transfer fails
Trigger: Threshold tuned on ECFP4 applied to RDKit FP, AtomPair, or MACCS.
Mechanism: Bit-density and fragment-resolution differ; Tanimoto distributions shift.
Symptom: "Similar" set is much larger or smaller than expected.
Fix: Re-tune the threshold per fingerprint and dataset. AtomPair ~0.55, MACCS ~0.85, ECFP4 ~0.7, and FCFP4 ~0.6 are repository starting heuristics, not universal equivalents.
Reconciliation: Cliffs Across Methods
If a pair flags as an activity cliff under one representation but not another, treat that as representation sensitivity. Inspect atom mappings, fingerprint environments, assay uncertainty, and the exact structural change; disagreement alone does not establish which substituent caused the activity difference.
Common Errors
Symptom
Cause
Fix
BulkTanimotoSimilarity output is treated as bit counts
The API returns similarity values, normally floats
Keep the returned values as similarities; inspect input vector types if the output is unexpected
Reported similarity > 1
Custom formula, malformed data, negative features, or an incorrectly normalized external implementation
Verify the coefficient definition and inputs; standard nonnegative RDKit Tanimoto and Tversky similarities are bounded by 1
Cluster centroids change when input order changes
Taylor-Butina tie handling and assignment are order-sensitive
Standardize and sort inputs by a stable identifier before clustering; record reordering and the input order