Use DELPHI Online
Commercially Available DELPHI No-Code Web Server
DELPHI: Interpretable Sequence-Based Antibody Developability Prediction for Polyreactivity and SEC Monomer Purity
DELPHI (Deep End-to-end Learning Platform for antibody developability with High Interpretability) is an open-source pipeline for training, benchmarking and interpreting sequence-based antibody developability classifiers. It was developed by Hoan N. Nguyen, Andre A. R. Teixeira and colleagues at the Institute for Protein Innovation in Boston, and described in a 2026 bioRxiv preprint. The authors state that the source code is released under the MIT license and that pre-trained polyreactivity and size-exclusion checkpoints are publicly available on Zenodo.
Many antibody candidates fail because of biophysical liabilities, such as polyreactivity (non-specific binding to unrelated proteins) and aggregation, and these are often found late and at high cost. Display technologies and next-generation sequencing can produce hundreds of thousands of binders per campaign, far more than can be screened experimentally. The authors note that developability remains an open prediction problem and that trained models and software are rarely shared, with most tools offering feature-level rather than residue-level explanations.
DELPHI addresses this by combining flexible model construction, leakage-aware training and interpretability in one tool. It pairs five classifier architectures with five antibody language models plus one-hot, biophysical and k-mer features, giving 25 ready-to-benchmark combinations.
Property | Detail |
|---|---|
Task | Binary developability classification (Pass or Fail) from antibody sequence |
Assays covered | Polyreactivity (PSR) and size-exclusion chromatography (SEC) monomericity |
Inputs | Antibody VH/VL or CDR H3 sequences with an assay label for training |
Classifiers | XGBoost, Random Forest, CNN, Transformer on language-model embeddings, dual-branch one-hot Transformer |
Validation | 10-fold CDR H3-cluster-stratified cross-validation, learning curves, external cohorts |
Availability | MIT-licensed code; pre-trained checkpoints on Zenodo |
How DELPHI Predicts Antibody Developability from Sequence
Label curation: assay readouts are converted into Pass or Fail labels, with class-imbalance correction.
CDR H3 clustering: sequences are clustered at 80% identity so that cross-validation folds do not leak near-identical CDR H3 loops between training and test.
Representation: antibody-specific language models (including AbLang2, IgBert, AntiBERTy and AntiBERTa2 variants) or one-hot, biophysical and k-mer features encode each sequence.
Classification: one of five architectures is trained, with hyperparameters that auto-scale to dataset size and class balance and remain configurable in YAML.
Thresholding and interpretation: a Youden-optimal operating threshold is stored in each checkpoint, and residue-level attributions (integrated gradients, SHAP) and in silico mutagenesis highlight risky positions.
DELPHI Benchmark Results for Polyreactivity and SEC Prediction
The authors trained on in-house yeast-display Fab data (11,265 antibodies for polyreactivity after balancing; 5,045 for SEC) and evaluated within and across libraries.
Evaluation | Result reported |
|---|---|
Within-library CV, polyreactivity (mean over 20 language-model combinations) | AUC 0.959 |
Within-library CV, SEC (mean over 20 combinations) | AUC 0.933 |
Best single combination, polyreactivity (XGBoost + IgBert) | AUC 0.967 |
Transfer to a public library of 246,293 antibodies | AUC up to 0.950 (AbLang2) |
Zero-shot, Jain clinical-stage panel (n = 137) | ROC-AUC 0.73 |
Zero-shot, Ginkgo PR-CHO | Spearman |rho| 0.35, similar to the best reported competition value of 0.337 |
Key Findings
Language model choice matters more than classifier: predictions from different architectures sharing one language model correlate strongly (Spearman 0.92 to 0.98).
Paired-chain models transfer best: AbLang2 and IgBert lead polyreactivity transfer across libraries.
Training-set size: performance plateaued at roughly 5,000 labelled antibodies in the polyreactivity analyses.
CDR H3 charge signature: arginine and lysine enrichment, aspartate depletion, CDR H3 length and net positive charge were associated with both polyreactivity and SEC failure.
What is Tamarind Bio?
Tamarind Bio is a no-code bioinformatics platform built to give life scientists and researchers access to powerful computational tools. Many cutting-edge machine learning models are hard to deploy and use. Tamarind provides an intuitive, web-based environment that removes the complexity of high-performance computing, software dependencies and command-line interfaces.
The platform is designed for biologists, chemists and other researchers who may not have a background in programming or cloud infrastructure but want to run models on their own data. Key features include:
A user-friendly graphical interface for setting up and launching experiments
A robust API for integration into existing research pipelines
An automated system for managing and scaling computational resources
Tamarind treats information and data security as a top priority, as detailed in its Trust Center and Terms of Service.
Accelerating Discovery with DELPHI on Tamarind Bio
Pre-screening antibody libraries: rank candidates from display or NGS campaigns for polyreactivity or monomer purity risk before running assays.
Residue-level engineering hypotheses: use attributions and CDR H3 mutagenesis to see which positions drive predicted risk, such as cationic residues in the loop.
Retraining for a new assay: benchmark language model and classifier combinations on your own labelled data and deploy the best one.
Multi-property triage: combine polyreactivity and SEC scores to prioritize leads that pass both.
How to Use DELPHI on Tamarind Bio
Log in and open the tool: sign in at tamarind.bio and select DELPHI from the tool list.
Provide antibody sequences: enter the heavy and light chain (or CDR H3) sequences of the antibodies to score, as the tool defines them.
Choose the model: select the polyreactivity or SEC predictor, if exposed; the paper deploys a Transformer with AbLang2 embeddings.
Set options: choose attribution or interpretation outputs, if exposed.
Run the job: submit it and wait for the predictions to complete.
Download results: retrieve the Pass or Fail probabilities and any residue-level attributions.
Interpret and iterate: compare scores with the stored operating threshold, then redesign flagged CDR H3 positions and rescore.
Things to Keep in Mind
The in-house training datasets are proprietary and cannot be shared; they come from a restricted germline space (8 VH and 6 VL germlines) with diversity confined to CDR H3.
Holding out entire VH germlines lowered the polyreactivity AUC from 0.965 to 0.903, so within-library numbers are optimistic for new germlines.
The polyreactivity Fail class is partly drawn from a different selection (NGS ssDNA) than the Pass class, so within-distribution AUC is best read as an upper estimate.
Predictions are binary classifications, not quantitative regression, and are computational hypotheses that need experimental confirmation.
Source: Nguyen HN, Kothiwal D, Su Y, Kieu MA, Cao R, Zhu H, Teixeira AAR. An interpretable open platform for sequence-based antibody developability prediction. bioRxiv preprint, 2026. DOI: 10.64898/2026.09.16.750421. Pre-trained checkpoints: 10.5281/zenodo.21823887.