Use CheMeleon Embeddings Online
Commercially Available CheMeleon Embeddings No-Code Web Server
CheMeleon Embeddings: Descriptor-Pretrained Molecular Foundation Model for Property Prediction and Similarity Search
CheMeleon is a molecular foundation model pre-trained on classical Mordred descriptors that produces learned molecular embeddings and strong starting weights for property prediction. It was developed by Jackson W. Burns, Akshat Shirish Zalte, William H. Green and colleagues at MIT and BASF SE, as described in the arXiv preprint (v2, 2026). The authors describe it as open source, permissively licensed and accessible through the Chemprop package.
Deep learning models for molecules, such as directed message-passing neural networks (D-MPNNs), often fail to beat classical methods like Random Forest on practical benchmarks with limited training data. Foundation models pre-trained on noisy experimental data or biased quantum mechanical simulations have not closed this gap reliably.
CheMeleon takes a different route: it pre-trains on deterministic, low-noise molecular descriptors, which compels the network to internalize expert chemical knowledge in a differentiable architecture.
Property | Detail |
|---|---|
Task | Molecular embeddings and fine-tuning for property prediction |
Architecture | D-MPNN in Chemprop; 2048-dimensional message passing, depth 6, mean aggregation |
Pre-training data | 1 million PubChem molecules with 1,613 Mordred descriptors |
Size | 8.7 million message-passing parameters; 12.9 million in total during pre-training |
Pre-training error | Test RMSE of 0.14 averaged across scaled descriptors |
Availability | Weights on Zenodo; integrated into Chemprop from version 2.2.0; code on GitHub |
How CheMeleon Learns Molecular Representations from Descriptors
Descriptor targets: Mordred descriptors are computed for one million randomly selected PubChem molecules, rescaled and Winsorized.
Pre-training: a D-MPNN with a feed-forward head is trained with a masked loss to predict the descriptors.
Transfer: the pre-trained message-passing weights are kept, the pre-trained head is discarded, and a new task-specific head is attached for fine-tuning.
Embeddings: the aggregated D-MPNN output serves as the molecular embedding for similarity analysis or downstream models.
CheMeleon Benchmark Results on Polaris and MoleculeACE
The authors evaluated CheMeleon on 58 benchmark datasets, counting a win when a model was best or statistically indistinguishable from the best.
Benchmark set | CheMeleon | Random Forest | Chemprop |
|---|---|---|---|
Polaris (win rate) | 75% | 68% | 32% |
MoleculeACE, entire test set (win rate) | 97% | 50% | 0% |
MoleculeACE, activity-cliff subset (win rate) | 100% | 67% | 13% |
Similarity and Toxicity Probing
In k-nearest-neighbor probing on 20 ToxCast endpoints, a read-across style test, fine-tuned CheMeleon embeddings reached the highest average balanced accuracy among the representations compared, with higher sensitivity at a modest cost in specificity. A Chemprop model of the same size trained from random initialization performed substantially worse, indicating that the gain comes from descriptor pre-training rather than model size.
What is Tamarind Bio?
Tamarind Bio is a no-code bioinformatics platform built to give life scientists and researchers access to powerful computational tools. Many cutting-edge machine learning models are hard to deploy and use. Tamarind provides an intuitive, web-based environment that removes the complexity of high-performance computing, software dependencies and command-line interfaces.
The platform is designed for biologists, chemists and other researchers who may not have a background in programming or cloud infrastructure but want to run models on their own data. Key features include:
A user-friendly graphical interface for setting up and launching experiments
A robust API for integration into existing research pipelines
An automated system for managing and scaling computational resources
Tamarind treats information and data security as a top priority, as detailed in its Trust Center and Terms of Service.
Accelerating Discovery with CheMeleon Embeddings on Tamarind Bio
Low-data property prediction: start from pre-trained weights for ADMET-style or assay datasets with limited labels.
Similarity and read-across: use embeddings to find chemically and biologically similar compounds, as in the toxicity probing study.
Activity-cliff-aware modeling: apply the model to datasets with structurally similar molecules of different potency.
Improving existing models: use the pre-trained D-MPNN to strengthen Chemprop-based workflows.
How to Use CheMeleon Embeddings on Tamarind Bio
Log in and open the tool: sign in at tamarind.bio and select CheMeleon Embeddings.
Provide molecules: enter or upload SMILES strings for the compounds of interest.
Choose the mode: generate embeddings from the pre-trained model, or fine-tune on labelled data, if exposed.
Set parameters: provide target columns and training options for fine-tuning, if exposed.
Run the job: submit and wait for completion.
Download outputs: retrieve the embedding vectors or predictions.
Use downstream: compute similarities, train simple models on the embeddings, or rank compounds.
Things to Keep in Mind
Many benchmarks had several statistically tied winners, so win rates do not imply uniform superiority.
Activity cliffs remain hard: CheMeleon was top and similar on both cliff and non-cliff compounds in only four cases, and the authors note the representation maps similar molecules together, which can hinder separating them.
The pre-training RMSE is reported for completeness; the model is not meant to compute descriptors, which Mordred calculates exactly.
The authors report that learned embeddings were obtained after fine-tuning on the training set for the kNN study.
Source: Burns JW, Zalte AS, Abreu CRA, Sieg J, Feldmann C, Mathea M, Green WH. Deep Learning Foundation Models from Classical Molecular Descriptors. arXiv:2506.15792v2 (2026). Weights: 10.5281/zenodo.15426600. Code: github.com/JacksonBurns/CheMeleon.