Use NDPredict Online
Commercially Available NDPredict No-Code Web Server
NDPredict: Structure-Based Prediction of Asparagine Deamidation Hotspots in Protein Therapeutics
NDPredict is a machine-learning approach for predicting asparagine (Asn) deamidation liabilities from protein 3D structure. It was developed by Lei Jia and Yaxiong Sun at Amgen and described in a PLOS ONE open-access article (2017). It combines experimentally measured penta-peptide deamidation half-lives with structural descriptors taken from crystal structures, and uses a random forest classifier to rank which Asn residues are most likely to deamidate.
Chemical stability is a major concern for protein therapeutics because it affects both efficacy and safety. Deamidation converts an Asn residue to aspartic acid or isoaspartic acid through a succinimide intermediate. If it happens in the complementarity determining region (CDR) of a monoclonal antibody, binding potency can be affected, so Asn deamidation liability is checked early during engineering.
In most pharmaceutical workflows this check is still sequence-based: motifs such as NG (and, to a lesser extent, NH, NS and NT) are flagged as liabilities. The authors note that this is convenient but less accurate, and often leads to over-engineering a protein by removing motifs that are not actually deamidated. NDPredict adds structural context to reduce those false alarms.
Property | Detail |
|---|---|
Task | Binary classification of Asn residues as deamidated or not |
Input | High-resolution protein structure with the Asn residue of interest |
Descriptors | 13 descriptors: penta-peptide half-life plus structure-based features |
Best model | Random forest (compared with SVM, naive Bayes, KNN, ANN and PLS) |
Training data | 194 Asn residues from 25 proteins (28 deamidated, 166 not) |
External test | 81 Asn residues from 3 proteins (5 deamidated, 76 not) |
How NDPredict Predicts Asn Deamidation from Structure
The model description in the paper rests on a descriptor set tailored to the deamidation mechanism:
Sequence-based half-life: the experimentally measured deamidation half-life of a Gly-Xxx-Asn-Yyy-Gly penta-peptide, which reflects the effect of the residue following Asn.
Nucleophilic attack distance: the distance between the Asn side-chain carbonyl carbon and the backbone nitrogen of the next residue, the basic geometric requirement for forming the succinimide intermediate.
Flexibility: crystallographic B-factors at the C, C-alpha, C-beta and C-gamma atoms, normalized as z-scores within each structure.
Solvent exposure: percent solvent accessibility of the residue and of its side chain.
Conformation: backbone phi and psi and side-chain chi1 and chi2 torsion angles, plus local secondary structure.
Feature importance
Recursive feature elimination on the random forest model ranked the penta-peptide half-life highest. Torsion angles and normalized B-factors also contributed strongly, with psi more important than phi and chi2 more important than chi1. The authors read this as evidence that backbone conformation influences deamidation more than side-chain conformation.
Benchmark Results: Accuracy of Structure-Based Deamidation Prediction
Models were cross-validated and then tested blind on the external set. Random forest had the best cross-validation accuracy (0.90) and the best enrichment on the blind test.
Blind-test metric | Random forest | SVM | Naive Bayes |
|---|---|---|---|
Accuracy | 0.95 | 0.94 | 0.75 |
AUC | 0.96 | 0.73 | 0.76 |
Recall | 0.80 | 0 | 0.60 |
Specificity | 0.96 | 1 | 0.76 |
Precision | 0.57 | - | 0.14 |
MCC | 0.65 | 0 | 0.20 |
The random forest model found 4 of the 5 deamidated residues in the test set with 3 false positives, while several other algorithms scored well on accuracy only by predicting every residue as non-deamidated. The paper also discusses why AUC, precision, recall and MCC are more informative than accuracy for unbalanced data like this.
What is Tamarind Bio?
Tamarind Bio is a no-code bioinformatics platform built to give life scientists and researchers access to powerful computational tools. Many cutting-edge machine learning models are hard to deploy and use. Tamarind provides an intuitive, web-based environment that removes the complexity of high-performance computing, software dependencies and command-line interfaces.
The platform is designed for biologists, chemists and other researchers who may not have a background in programming or cloud infrastructure but want to run models on their own data. Key features include:
A user-friendly graphical interface for setting up and launching experiments
A robust API for integration into existing research pipelines
An automated system for managing and scaling computational resources
Tamarind treats information and data security as a top priority, as detailed in its Trust Center and Terms of Service.
Accelerating Discovery with NDPredict on Tamarind Bio
Antibody and biologic developability: screen Asn residues, including those in CDRs, for deamidation risk before expression and formulation work.
Avoiding over-engineering: check whether an NG or other liable motif is actually predicted to deamidate in its structural context before mutating it.
Ranking hotspots: prioritize which Asn residues deserve experimental stability testing first.
Other hotspots: the authors suggest similar structure-function models may extend to other chemical liabilities.
How to Use NDPredict on Tamarind Bio
Open the tool: log in to tamarind.bio and select NDPredict from the tool list.
Provide a structure: upload a protein structure (for example an experimental or predicted model of your protein or antibody), as the interface requests.
Select residues: if exposed, choose the chains or Asn residues to evaluate; otherwise all Asn residues are scored.
Run the job: submit the run and wait for the prediction to finish.
Download results: retrieve the per-residue deamidation predictions.
Interpret: compare flagged residues with sequence motifs such as NG, and focus on residues the model ranks highest.
Use downstream: guide mutagenesis, developability review or experimental forced-degradation studies.
Things to Keep in Mind
The training set is small (194 Asn residues from 25 proteins) and the external test set has only 5 deamidated residues, so metrics carry uncertainty.
The random forest correctly identified only 4 of 7 non-deamidated NG motifs in the test set, which the authors call acceptable but not very high.
B-factors are sensitive to resolution, so the method was built on high-resolution crystal structures; results on lower-quality or modeled structures may differ.
The model gives a binary deamidated or not call; the authors suggest richer descriptors and more data as future improvements.
Source: Jia L, Sun Y. Protein asparagine deamidation prediction based on structures with machine learning methods. PLOS ONE 12(7): e0181347 (2017). DOI: https://doi.org/10.1371/journal.pone.0181347. The paper reports no code repository; training and test data are in its supporting information.