Levenshtein similarity

Molecule Activity Cliff Estimation (MoleculeACE) is a tool for evaluating the predictive performance on activity cliff compounds of machine learning models.

MoleculeACE can be used to:

Analyze and compare the performance on activity cliffs of machine learning methods typically employed in QSAR.
Identify best practices to enhance a model’s predictivity in the presence of activity cliffs.
Design guidelines to consider when developing novel QSAR approaches.

Update:

Upon request, we added an extra column to the datasets containing pEC50 and pKi values calculated from Molar concentrations alongside the original training labels used in the study that used log-transformed nM concentrations. Model errors will be the same when trained with either log transformed nM or log transformed M values (except for random processes), since labels are simple shiften by 9.

📖 Table of Contents

Table of Contents

➤ Benchmark study
➤ Tool
➤ Prerequisites
➤ Installation
- Pip installation
- Manual installation
➤ Getting started
- Train an out-of-the-box model
- Evaluate your own model
➤ How to cite
➤ Licence

Benchmark study

In a benchmark study we collected and curated bioactivity data on 30 macromolecular targets, which were used to evaluate the performance of many machine learning algorithms on activity cliffs. We used classical machine learning methods combined with common molecular descriptors and neural networks based on unstructured molecular data like molecular graphs or SMILES strings.

Activity cliffs are molecules with small differences in structure but large differences in potency. Activity cliffs play an important role in drug discovery, but the bioactivity of activity cliff compounds are notoriously difficult to predict.

Example of an activity cliff on the Dopamine D3 receptor, D3R

Tool

Any regression model can be evaluated on activity cliff performance using MoleculeACE on third party data or the 30 included molecular bioactivity data sets. All 24 machine learning strategies covered in our benchmark study can be used out of the box.

Prerequisites

MoleculeACE currently supports Python 3.8. Some required deep learning packages are not included in the pip install.

Tensorflow (2.9.0)
PyTorch (1.11.0)
PyTorch Geometric (2.0.4)
Transformers (4.20.1)

Installation

Pip installation

MoleculeACE can be installed as

pip install MoleculeACE

Manual installation

git clone https://github.com/molML/MoleculeACE.git

pip install rdkit-pypi pandas numpy pandas chembl_webresource_client scikit-learn matplotlib tqdm python-Levenshtein

Getting started

Train an out-of-the-box model on one of the many included datasets

from MoleculeACE import MPNN, Data, Descriptors, calc_rmse, calc_cliff_rmse, get_benchmark_config

dataset = 'CHEMBL2034_Ki'
descriptor = Descriptors.GRAPH
algorithm = MPNN

# Load data
data = Data(dataset)

# Get the already optimized hyperparameters
hyperparameters = get_benchmark_config(dataset, algorithm, descriptor)

# Featurize SMILES strings with a specific method
data(descriptor)

# Train and a model
model = algorithm(**hyperparameters)
model.train(data.x_train, data.y_train)
y_hat = model.predict(data.x_test)

# Evaluate your model on activity cliff compounds
rmse = calc_rmse(data.y_test, y_hat)
rmse_cliff = calc_cliff_rmse(y_test_pred=y_hat, y_test=data.y_test, cliff_mols_test=data.cliff_mols_test)

print(f"rmse: {rmse}")
print(f"rmse_cliff: {rmse_cliff}")

Evaluate the performance of your own model

from MoleculeACE import calc_rmse, calc_cliff_rmse

# Train your own model
model = ...
y_hat = model.predict(...)

# Evaluate your model on activity cliff compounds
rmse = calc_rmse(y_test, y_hat)
# You need to provide both the predicted and true values of the test set + train labels + the train and test molecules
# Activity cliffs are calculated on the fly
rmse_cliff = calc_cliff_rmse(y_test_pred=y_hat, y_test=y_test, smiles_test=smiles_test, y_train=y_train, 
                             smiles_train=smiles_train, in_log10=True, similarity=0.9, potency_fold=10)

print(f"rmse: {rmse}")
print(f"rmse_cliff: {rmse_cliff}")

How to cite

Exposing the Limitations of Molecular Machine Learning with Activity Cliffs. Derek van Tilborg, Alisa Alenicheva, and Francesca Grisoni. Journal of Chemical Information and Modeling, 2022, 62 (23), 5938-5951. DOI: 10.1021/acs.jcim.2c01073

License

MoleculeACE is under MIT license. For use of specific models, please refer to the model licenses found in the original packages.

smiles	exp_mean [nM]	y
Cc1cncc(-c2cc3c(-c4cccc(N5CCNCC5)n4)n[nH]c3cn2)n1	100	2
Cc1ccc(F)c(-c2nc(C(=O)Nc3cnn(C)c3N3CCCC@@HCC3)c(N)s2)c1F	100	2
Cn1ncc(NC(=O)c2nc(-c3ccccc3F)sc2N)c1N1CCC@HCC(F)(F)C1	100	2
Nc1sc(-c2c(F)cccc2F)nc1C(=O)Nc1cnn(C2CC2)c1N1CCC@HCC(F)(F)C1	100	2
C=C(C)c1ccc(-c2n[nH]c3cnc(-c4cccnc4)cc23)nc1N1CCCC@HC1	100	2
C#Cc1ccc(-c2n[nH]c3cnc(-c4cccnc4)cc23)nc1N1CCCC@HC1	100	2
Cn1ncc(NC(=O)c2nc(-c3ccc(C(F)(F)F)cc3F)sc2N)c1[C@@h]1CCC@@H C@@HCO1	100	2
CO[C@H]1COC@HCC[C@H]1N	100	2
Cn1ncc(NC(=O)c2csc(-c3c(F)cc(C4(F)COC4)cc3F)n2)c1[C@@h]1CCC@@H C@HCO1	100	2
Nc1sc(-c2c(F)cccc2F)nc1C(=O)Nc1cnccc1N1CCCC@HC1	100	2
CN1CCC(N(C)c2ccc3nnc(-c4cccc(C(F)(F)F)c4)n3n2)CC1	100	-2
Cn1c2ccccc2c2c3c(c4c5ccccc5n(CCC#N)c4c21)CNC3=O	100	-2
c1ccc(CNc2cc(-c3c[nH]c4ncccc34)ncn2)cc1	100	-2
CSc1ccc2nc3c(c(Cl)c2c1)CCNC3=O	100	-2
Cc1n[nH]c2ccc(-c3cncc(OCC(N)Cc4ccccc4)c3)cc12	100	-2
O=c1[nH]c2sc3c(c2c2nc(-c4ccccc4)nn12)CCCC3	100	-2
O=C1NC(=O)C(c2c[nH]c3ccccc23)=C1c1nc(N2CCNCC2)nc2ccccc12	100	-2

molml / moleculeace Goto Github PK

moleculeace's Introduction

Update:

📖 Table of Contents

Benchmark study

Tool

Prerequisites

Installation

Pip installation

Manual installation

Getting started

Train an out-of-the-box model on one of the many included datasets

Evaluate the performance of your own model

How to cite

License

moleculeace's People

Contributors

Stargazers

Watchers

Forkers

moleculeace's Issues

Recommend Projects

Recommend Topics

Recommend Org