Use NDPredict Online

Commercially Available NDPredict No-Code Web Server

NDPredict: Structure-Based Prediction of Asparagine Deamidation Hotspots in Protein Therapeutics

NDPredict is a machine-learning approach for predicting asparagine (Asn) deamidation liabilities from protein 3D structure. It was developed by Lei Jia and Yaxiong Sun at Amgen and described in a PLOS ONE open-access article (2017). It combines experimentally measured penta-peptide deamidation half-lives with structural descriptors taken from crystal structures, and uses a random forest classifier to rank which Asn residues are most likely to deamidate.

Chemical stability is a major concern for protein therapeutics because it affects both efficacy and safety. Deamidation converts an Asn residue to aspartic acid or isoaspartic acid through a succinimide intermediate. If it happens in the complementarity determining region (CDR) of a monoclonal antibody, binding potency can be affected, so Asn deamidation liability is checked early during engineering.

In most pharmaceutical workflows this check is still sequence-based: motifs such as NG (and, to a lesser extent, NH, NS and NT) are flagged as liabilities. The authors note that this is convenient but less accurate, and often leads to over-engineering a protein by removing motifs that are not actually deamidated. NDPredict adds structural context to reduce those false alarms.

Property

Detail

Task

Binary classification of Asn residues as deamidated or not

Input

High-resolution protein structure with the Asn residue of interest

Descriptors

13 descriptors: penta-peptide half-life plus structure-based features

Best model

Random forest (compared with SVM, naive Bayes, KNN, ANN and PLS)

Training data

194 Asn residues from 25 proteins (28 deamidated, 166 not)

External test

81 Asn residues from 3 proteins (5 deamidated, 76 not)

How NDPredict Predicts Asn Deamidation from Structure

The model description in the paper rests on a descriptor set tailored to the deamidation mechanism:

  1. Sequence-based half-life: the experimentally measured deamidation half-life of a Gly-Xxx-Asn-Yyy-Gly penta-peptide, which reflects the effect of the residue following Asn.

  2. Nucleophilic attack distance: the distance between the Asn side-chain carbonyl carbon and the backbone nitrogen of the next residue, the basic geometric requirement for forming the succinimide intermediate.

  3. Flexibility: crystallographic B-factors at the C, C-alpha, C-beta and C-gamma atoms, normalized as z-scores within each structure.

  4. Solvent exposure: percent solvent accessibility of the residue and of its side chain.

  5. Conformation: backbone phi and psi and side-chain chi1 and chi2 torsion angles, plus local secondary structure.

Feature importance

Recursive feature elimination on the random forest model ranked the penta-peptide half-life highest. Torsion angles and normalized B-factors also contributed strongly, with psi more important than phi and chi2 more important than chi1. The authors read this as evidence that backbone conformation influences deamidation more than side-chain conformation.

Benchmark Results: Accuracy of Structure-Based Deamidation Prediction

Models were cross-validated and then tested blind on the external set. Random forest had the best cross-validation accuracy (0.90) and the best enrichment on the blind test.

Blind-test metric

Random forest

SVM

Naive Bayes

Accuracy

0.95

0.94

0.75

AUC

0.96

0.73

0.76

Recall

0.80

0

0.60

Specificity

0.96

1

0.76

Precision

0.57

-

0.14

MCC

0.65

0

0.20

The random forest model found 4 of the 5 deamidated residues in the test set with 3 false positives, while several other algorithms scored well on accuracy only by predicting every residue as non-deamidated. The paper also discusses why AUC, precision, recall and MCC are more informative than accuracy for unbalanced data like this.

What is Tamarind Bio?

Tamarind Bio is a no-code bioinformatics platform built to give life scientists and researchers access to powerful computational tools. Many cutting-edge machine learning models are hard to deploy and use. Tamarind provides an intuitive, web-based environment that removes the complexity of high-performance computing, software dependencies and command-line interfaces.

The platform is designed for biologists, chemists and other researchers who may not have a background in programming or cloud infrastructure but want to run models on their own data. Key features include:

  • A user-friendly graphical interface for setting up and launching experiments

  • A robust API for integration into existing research pipelines

  • An automated system for managing and scaling computational resources

Tamarind treats information and data security as a top priority, as detailed in its Trust Center and Terms of Service.

Accelerating Discovery with NDPredict on Tamarind Bio

  • Antibody and biologic developability: screen Asn residues, including those in CDRs, for deamidation risk before expression and formulation work.

  • Avoiding over-engineering: check whether an NG or other liable motif is actually predicted to deamidate in its structural context before mutating it.

  • Ranking hotspots: prioritize which Asn residues deserve experimental stability testing first.

  • Other hotspots: the authors suggest similar structure-function models may extend to other chemical liabilities.

How to Use NDPredict on Tamarind Bio

  1. Open the tool: log in to tamarind.bio and select NDPredict from the tool list.

  2. Provide a structure: upload a protein structure (for example an experimental or predicted model of your protein or antibody), as the interface requests.

  3. Select residues: if exposed, choose the chains or Asn residues to evaluate; otherwise all Asn residues are scored.

  4. Run the job: submit the run and wait for the prediction to finish.

  5. Download results: retrieve the per-residue deamidation predictions.

  6. Interpret: compare flagged residues with sequence motifs such as NG, and focus on residues the model ranks highest.

  7. Use downstream: guide mutagenesis, developability review or experimental forced-degradation studies.

Things to Keep in Mind

  • The training set is small (194 Asn residues from 25 proteins) and the external test set has only 5 deamidated residues, so metrics carry uncertainty.

  • The random forest correctly identified only 4 of 7 non-deamidated NG motifs in the test set, which the authors call acceptable but not very high.

  • B-factors are sensitive to resolution, so the method was built on high-resolution crystal structures; results on lower-quality or modeled structures may differ.

  • The model gives a binary deamidated or not call; the authors suggest richer descriptors and more data as future improvements.

Source: Jia L, Sun Y. Protein asparagine deamidation prediction based on structures with machine learning methods. PLOS ONE 12(7): e0181347 (2017). DOI: https://doi.org/10.1371/journal.pone.0181347. The paper reports no code repository; training and test data are in its supporting information.

Supporting 10,000+ scientists around the world,

from leading biotechs, and global biopharma