Use Chemprop Online
Commercially Available Chemprop No-Code Web Server
Chemprop: Directed Message-Passing Neural Networks for Molecular Property Prediction
Chemprop is an open-source machine learning package that implements the directed message-passing neural network (D-MPNN) for chemical property prediction. The 2024 Journal of Chemical Information and Modeling application note, by Esther Heid, Kevin P. Greenman, Yunsie Chung, Shih-Cheng Li, David E. Graff, Florence H. Vermeire, Haoyang Wu, William H. Green and Charles J. McGill, describes its current feature set. Chemprop is available under the MIT licence on GitHub.
Deep learning is now widely used to predict molecular properties, and many representations and architectures compete: graph networks, transformers on SMILES or SELFIES, and 3D-aware models. Chemprop focuses on graph convolution over the molecular graph, which performs robustly for properties that depend on local structure and when no 3D conformation is known or relevant. Its aim is simple, fast access to learned molecular properties for non-experts.
Property | Detail |
|---|---|
Architecture | Directed message-passing neural network (D-MPNN) followed by a feed-forward network |
Inputs | Single molecules, multi-molecule systems (for example solute/solvent), atom-mapped reactions |
Targets | Molecular properties, atom/bond-level properties, spectra |
Uncertainty | Ensembling, mean-variance estimation, evidential learning, plus calibration methods |
Availability | MIT licence; github.com/chemprop/chemprop with documentation and tutorials |
How Chemprop Predicts Molecular Properties with a D-MPNN
Molecular graph: atoms are nodes with features such as chirality, hydrogen count, hybridisation, aromaticity and mass; bonds are edges with type, conjugation, ring and stereo features.
Directed message passing: each bond has two directed edges, and information is passed along them to update the local neighbourhood.
Aggregation: atom representations are aggregated into a molecule-level representation.
Feed-forward prediction: a feed-forward network maps this latent representation to the target. The latent vectors can also be output for embedding analysis.
New Chemprop Capabilities: Reactions, Spectra, Atom-Level Targets and Uncertainty
Reactions and Multi-Molecule Models
Chemprop accepts atom-mapped reactions, represents them as a condensed graph of reaction (CGR), and can include solvent. Multi-molecule models embed each molecule and concatenate the embeddings.
Spectra and Atom/Bond-Level Targets
Spectral targets such as IR absorbance are supported, with spectral information divergence as a loss, handling of missing values and optional exclusion regions. A multitask constrained D-MPNN predicts atom and bond properties, optionally constraining them to sum to a molecular value such as net charge.
Training Workflow Tools
Uncertainty quantification: estimation, calibration and evaluation metrics for regression, classification and spectra.
Transfer learning: initialise from a pretrained model or freeze D-MPNN and selected feed-forward layers.
Hyperparameter optimisation: improved search over trials.
Customisation: loss functions and atom/bond features.
The authors report state-of-the-art performance on water-octanol partition coefficients, reaction barrier heights, atomic partial charges and absorption spectra; detailed benchmarks are in the paper's Supporting Information.
What is Tamarind Bio?
Tamarind Bio is a no-code bioinformatics platform built to give life scientists and researchers access to powerful computational tools. Many cutting-edge machine learning models are hard to deploy and use. Tamarind provides an intuitive, web-based environment that removes the complexity of high-performance computing, software dependencies and command-line interfaces.
The platform is designed for biologists, chemists and other researchers who may not have a background in programming or cloud infrastructure but want to run models on their own data. Key features include:
A user-friendly graphical interface for setting up and launching experiments
A robust API for integration into existing research pipelines
An automated system for managing and scaling computational resources
Tamarind treats information and data security as a top priority, as detailed in its Trust Center and Terms of Service.
Accelerating Discovery with Chemprop on Tamarind Bio
Physicochemical and pharmacological properties: train models on your own labelled molecules, such as lipophilicity or solvation.
Reaction properties: predict quantities such as reaction barrier heights from atom-mapped reactions.
Spectra: predict absorption or IR spectra for chemical analysis.
Atom-level properties: predict quantities such as partial charges.
How to Use Chemprop on Tamarind Bio
Open the tool: log in to tamarind.bio and select Chemprop.
Provide data: supply SMILES with target values; use atom-mapped reaction SMILES for reactions or several SMILES columns for multi-molecule systems, if exposed.
Choose the task: regression, classification, multiclass or spectra, if exposed.
Set parameters: options such as loss function, uncertainty method, or transfer from a pretrained model, if exposed.
Run: train and evaluate the model.
Download outputs: collect predictions, uncertainty estimates and trained models.
Use downstream: prioritise or filter candidate molecules by predicted properties.
Things to Keep in Mind
The D-MPNN operates on the 2D molecular graph and is best suited to properties driven by local structure, not to properties that depend on 3D conformation.
Reaction inputs must be atom-mapped.
Model quality depends on the quantity and quality of training data; detailed benchmarks are in the Supporting Information.
Source: Heid E, Greenman KP, Chung Y, Li S-C, Graff DE, Vermeire FH, Wu H, Green WH, McGill CJ. Chemprop: A Machine Learning Package for Chemical Property Prediction. J. Chem. Inf. Model. 2024, 64, 9-17. DOI: 10.1021/acs.jcim.3c01250. Code: https://github.com/chemprop/chemprop.