Use HELM-GPT Online
Commercially Available HELM-GPT No-Code Web Server
HELM-GPT: De Novo Macrocyclic Peptide Design with a Generative Pre-Trained Transformer and Reinforcement Learning
HELM-GPT is a generative method for de novo macrocyclic peptide design that combines the Hierarchical Editing Language for Macromolecules (HELM) representation with a generative pre-trained transformer (GPT). It was developed by Xiaopeng Xu, Xin Gao and colleagues at King Abdullah University of Science and Technology (KAUST) and Syneron Technology, and published in Bioinformatics in 2024. The authors state that the code and data are freely available on GitHub.
Macrocyclic peptides can bind flat protein surfaces with high affinity and specificity and may cross the cell membrane, which makes them attractive for intracellular targets such as KRAS. Computational approaches for de novo macrocyclic peptide design, however, remain largely unexplored. Generative models built on SMILES strings or molecular graphs often produce molecules that are hard to synthesize, while plain amino acid sequences cannot describe the non-natural amino acids and cyclizing bonds that these peptides contain.
HELM-GPT works in HELM space with a library of predefined synthesizable monomers, so generated peptides are built from monomers that can actually be incorporated.
Property | Detail |
|---|---|
Task | Generate and optimize macrocyclic peptides, including those with non-natural amino acids |
Representation | HELM sequences, converted to SMILES with RDKit for property prediction |
Model | Transformer decoder with eight decoder blocks, trained by next-token prediction |
Optimization | Reinforcement learning with the Reinvent loss plus a contrastive preference learning (CPL) loss |
Monomer library | 3,104 monomers from ChEMBL, CycPeptMPDB and KRAS HELMs |
Availability | Code and data on GitHub |
How HELM-GPT Generates and Optimizes Macrocyclic Peptides
Pretraining: a GPT prior is pretrained on 22,040 HELM sequences from ChEMBL.
Fine-tuning: the prior is fine-tuned on 7,451 cyclic peptides from CycPeptMPDB, optionally with 226 KRAS-related peptides from patents.
Property predictors: a random forest model predicts membrane permeability (Spearman 0.82 on test data) and an XGBoost model predicts KRAS binding Kd (Spearman 0.82).
Reinforcement learning: an agent generates batches of HELM sequences, scores them with the predictors, and updates its likelihoods using the Reinvent loss and the CPL loss, which compares pairs of molecules by preference.
Step-by-step co-optimization: permeability is optimized first, passing molecules are used to fine-tune a new prior, and both properties are then optimized together.
HELM-GPT Results: Permeability, KRAS Binding and Peptide Validity
Evaluation | Result reported |
|---|---|
Validity of the ChEMBL-pretrained prior | 70.8% valid, 89.0% unique, 88.9% novel |
Validity after fine-tuning on CycPeptMPDB | 83.9% valid, 91.3% unique |
Validity after fine-tuning on KRAS plus CycPeptMPDB | 100% valid, 100% unique |
Mean predicted permeability, HELM-GPT (Reinvent + CPL) | -4.414, versus -4.589 for the best compared baseline (SMILES LSTM hill climbing) |
Predicted KRAS Kd score, HELM-GPT (Reinvent + CPL) | 0.038, second to Graph GA (-1.855) |
Molecules passing both filters, direct co-optimization | 20 |
Molecules passing both filters, step-by-step strategy | 17,273 (a 785-fold improvement) |
Key Findings
CPL helps: adding the contrastive preference loss generally improved optimization over the Reinvent loss alone, except for diversity.
Fast optimization: predicted permeability rose within about 150 RL steps and predicted KRAS Kd improved within about 100 steps.
Synthesizability by design: restricting generation to predefined monomers keeps outputs within a synthesizable HELM space.
What is Tamarind Bio?
Tamarind Bio is a no-code bioinformatics platform built to give life scientists and researchers access to powerful computational tools. Many cutting-edge machine learning models are hard to deploy and use. Tamarind provides an intuitive, web-based environment that removes the complexity of high-performance computing, software dependencies and command-line interfaces.
The platform is designed for biologists, chemists and other researchers who may not have a background in programming or cloud infrastructure but want to run models on their own data. Key features include:
A user-friendly graphical interface for setting up and launching experiments
A robust API for integration into existing research pipelines
An automated system for managing and scaling computational resources
Tamarind treats information and data security as a top priority, as detailed in its Trust Center and Terms of Service.
Accelerating Discovery with HELM-GPT on Tamarind Bio
Targeting intracellular proteins: propose macrocyclic peptides for targets without drug-like pockets, such as KRAS.
Cell-penetrating peptides: generate cyclic peptides with high predicted membrane permeability.
Multi-objective design: co-optimize permeability and binding affinity with the step-by-step strategy.
Library design: build large sets of filter-passing candidates to prioritize for experimental screening.
How to Use HELM-GPT on Tamarind Bio
Log in and open the tool: sign in at tamarind.bio and select HELM-GPT.
Choose the design objective: pick a generation or optimization task such as permeability or KRAS binding, if exposed.
Set generation parameters: specify the number of peptides to sample and any optimization settings the tool exposes.
Run the job: submit it and wait for the HELM sequences to be generated.
Download outputs: retrieve the generated HELM sequences and their predicted property scores.
Filter candidates: apply thresholds like those in the paper (permeability above -6.0, KRAS Kd below 10) to shortlist peptides.
Move downstream: evaluate the shortlist with structure prediction or docking, then synthesize and test the best candidates.
Things to Keep in Mind
Reported results rely on predicted properties; the authors note that predicted scores cannot guarantee good experimental properties.
Property predictors must be reliable, otherwise the agent can generate many false positives.
Generation is limited to the predefined monomer set, which must be chosen to suit each project.
Co-optimizing permeability and KRAS binding directly was difficult, which motivated the step-by-step strategy.
HELM-GPT was less diverse than SMILES- and graph-based genetic algorithms because it is constrained to HELM space.
Source: Xu X, Xu C, He W, Wei L, Li H, Zhou J, Zhang R, Wang Y, Xiong Y, Gao X. HELM-GPT: de novo macrocyclic peptide design using generative pre-trained transformer. Bioinformatics 2024;40(6):btae364. DOI: 10.1093/bioinformatics/btae364. Code: github.com/charlesxu90/helm-gpt.