Use FoldBench Evaluator Online

Commercially Available FoldBench Evaluator No-Code Web Server

FoldBench Evaluator: Benchmarking All-Atom Biomolecular Structure Prediction Across Proteins, Nucleic Acids and Ligands

FoldBench is a unified, low-homology benchmark for all-atom biomolecular structure prediction, built to compare models such as AlphaFold 3, Boltz-1, Chai-1, HelixFold 3 and Protenix on the same targets and metrics. It was introduced by Sheng Xu, Qiantai Feng, Siqi Sun and colleagues, and published in Nature Communications. The authors release the benchmark data, evaluation code and reference results.

All-atom predictors now cover proteins, nucleic acids, ligands and ions, but until FoldBench there was no cross-domain benchmark to compare them. Earlier studies usually addressed a single task, covered few methods, and often ignored how accuracy changes with similarity to the training data. Without unified dataset construction and evaluation criteria, published numbers are hard to compare and may not reflect real generalization.

Property

Detail

Benchmark size

1522 biological assemblies across nine prediction tasks

Tasks

Protein, RNA and DNA monomers; protein-protein, antibody-antigen, protein-ligand, protein-RNA, protein-DNA and protein-peptide interfaces

Data window

PDB entries released after 2023-01-13 and before 2024-11-01, filtered for low homology to training data

Metrics

DockQ, ligand RMSD with LDDT-PLI, LDDT, computed with OpenStructure

Models evaluated

AlphaFold 3, Boltz-1, Chai-1, HelixFold 3, Protenix

Why a Low-Homology All-Atom Structure Prediction Benchmark Matters

Structure prediction models can look strong when test complexes resemble their training data. FoldBench removes targets with high sequence or structural similarity to the training set, so scores speak to generalization rather than memorization. The final set holds 334 protein monomers, 15 RNA monomers and 14 DNA monomers, plus 279 protein-protein, 172 antibody-antigen, 558 protein-ligand, 70 protein-RNA, 330 protein-DNA and 51 protein-peptide interfaces.

How the FoldBench Evaluation Works

  1. Curate targets: bioassemblies from the PDB are filtered by sequence identity (40 percent for proteins) and ligand similarity (0.5 Tanimoto), then clustered, keeping the highest-resolution structure per cluster.

  2. Generate predictions: each model is run with a 5 x 5 sampling strategy (5 seeds, 5 samples) and 10 recycles.

  3. Score interfaces: DockQ is used for most protein interfaces, where a score of at least 0.23 counts as a success. Ligand success means ligand RMSD under 2 angstrom and LDDT-PLI above 0.8.

  4. Stratify by similarity: results are split into subsets such as unseen protein and unseen ligand, so accuracy can be tracked against similarity to training data.

Key Results from the FoldBench Benchmark

AlphaFold 3 gave the best accuracy across most tasks, but the study highlights clear weak spots for every model.

Finding

Reported result

Protein-ligand success, all 558 targets

AlphaFold 3 at 64.9 percent, nearly 10 points above runner-up Boltz-1

Unseen protein subset (76 targets)

AlphaFold 3 at 69.0 percent

Unseen ligand subset (482 targets)

AlphaFold 3 at 64.3 percent

Protein-protein DockQ success

AlphaFold 3 at 72.9 percent

Antibody-antigen

Failure rates above 50 percent for current methods

Ligand and Protein Interface Generalization

Ligand docking accuracy falls as ligand similarity to the training set decreases, and the authors conclude that ligand similarity, more than protein modeling quality, limits protein-ligand accuracy. A similar pattern appears for protein-protein interfaces.

Antibody-Antigen and Nucleic Acid Challenges

Antibody-antigen prediction remains hard because CDR loops are flexible and confidence scores do not reliably pick near-native poses. Nucleic acids are also difficult, with the paper reporting average LDDT of 0.53 for DNA and 0.61 for RNA for the best model in those tasks.

What is Tamarind Bio?

Tamarind Bio is a no-code bioinformatics platform built to give life scientists and researchers access to powerful computational tools. Many cutting-edge machine learning models are hard to deploy and use. Tamarind provides an intuitive, web-based environment that removes the complexity of high-performance computing, software dependencies and command-line interfaces.

The platform is designed for biologists, chemists and other researchers who may not have a background in programming or cloud infrastructure but want to run models on their own data. Key features include:

  • A user-friendly graphical interface for setting up and launching experiments

  • A robust API for integration into existing research pipelines

  • An automated system for managing and scaling computational resources

Tamarind treats information and data security as a top priority, as detailed in its Trust Center and Terms of Service.

Accelerating Discovery with FoldBench on Tamarind Bio

  • Model selection: compare all-atom predictors on a common footing before choosing one for a project.

  • Ligand and pocket confidence: judge how far to trust co-folding predictions for ligands unlike known training complexes.

  • Antibody and nucleic acid projects: understand expected failure modes before relying on predicted complexes.

  • Method development: test a new structure prediction model against the same low-homology targets and metrics.

How to Use FoldBench Evaluator on Tamarind Bio

  1. Log in and open the tool: sign in at tamarind.bio and select FoldBench Evaluator.

  2. Provide predictions: supply predicted structures for the benchmark targets (if exposed), in the format the tool asks for.

  3. Choose the task: select the interface or monomer category to evaluate, such as protein-ligand or antibody-antigen (if exposed).

  4. Run the evaluation: submit the job so predictions are scored against the reference structures.

  5. Download results: retrieve the per-target scores and summary tables.

  6. Interpret: compare success rates by subset, using the DockQ 0.23 threshold and the ligand RMSD plus LDDT-PLI criterion defined in the paper.

Options on Tamarind may differ from the full benchmark repository, so check the tool page for what is exposed.

Things to Keep in Mind

  • The benchmark covers PDB entries from 2023-01-13 to 2024-11-01, so it reflects that window only.

  • Targets are low-homology by design, and results should not be read as accuracy on well-represented complexes.

  • Some models could not process every input: for example Boltz-1 cannot handle glycans and the evaluated Chai-1 version could not handle certain targets, so those were handled specially.

  • Benchmark scores describe prediction accuracy against experimental structures, not binding affinity or function.

Source: Xu, Feng, Qiao, Wu, Shen, Cheng, Zheng and Sun, "Benchmarking all-atom biomolecular structure prediction with FoldBench," Nature Communications (2026) 17:442, DOI 10.1038/s41467-025-67127-3. Code and data: github.com/BEAM-Labs/FoldBench.

Supporting 10,000+ scientists around the world,

from leading biotechs, and global biopharma