Use MATCHA Online

Commercially Available MATCHA No-Code Web Server

MATCHA: Multi-Stage Riemannian Flow Matching for Fast, Physically Valid Protein-Ligand Molecular Docking

MATCHA (Multi-stage Riemannian flow matching for accurate and physically valid molecular docking) is a protein-ligand docking pipeline from Daria Frolova, Talgat Daulbaev and colleagues at Ligand Pro and the Skolkovo Institute of Science and Technology. It combines three sequential flow matching stages with GNINA energy minimization and physical validity filters. Inference code and model weights are publicly available on GitHub.

Predicting ligand binding poses is central to structure-based drug design and virtual screening, but existing methods struggle to balance speed, accuracy and physical plausibility. Large co-folding models are accurate but slow, and many deep learning docking methods return poses with clashes or bad geometry. Benchmarks also differ widely in targets and ligands, which makes generalization hard to judge.

Property

Detail

Task

Protein-ligand docking, blind by default, with a pocket-aware setting

Approach

Three flow matching stages over translation, rotation and torsions, then GNINA minimization and PoseBusters-style filters

Ligand model

Semi-flexible ligand, rigid protein

Protein features

ESM-2 35M embeddings

Training data

PDBBind and Binding MOAD

Variants

MATCHA and a faster MATCHA-LITE

How MATCHA Predicts Protein-Ligand Binding Poses

  1. Stage 1, coarse placement: starting from a random initialization, flow matching integrates translation, rotation and torsion degrees of freedom to place the ligand.

  2. Stage 2, refinement: starting from the predicted translation with uniformly distributed angles, a second model refines orientation and torsions.

  3. Stage 3, final polish: a third model with small noise produces the final refined pose.

  4. GNINA minimization: candidate poses from all stages are locally minimized and ranked by GNINA affinity.

  5. Physical validity filtering: a minimal set of unsupervised PoseBusters checks removes unrealistic complexes, and the best remaining pose is selected.

MATCHA Docking Accuracy, Physical Validity and Speed

Benchmark

Reported result

Astex Diverse Set (n = 85)

85.9 percent at RMSD 2 angstrom or less; 82.4 percent when also PoseBusters-valid, 11.8 points above AlphaFold 3

PDBBind test set (n = 363)

49.6 percent RMSD 2 angstrom or less and PB-valid, versus 43.8 percent for AlphaFold 3

PoseBusters V2 held-out subset (n = 130)

Stable performance while co-folding methods drop by up to about 10 percent

Speed

About 31 times faster than AlphaFold 3, Chai-1 and Boltz-2; about 12.7 s per complex, or about 7.7 s for MATCHA-LITE, on an A100 40GB

Physical Plausibility

On PoseBusters V2, 95.6 percent of MATCHA poses within 2 angstrom RMSD are physically valid. The authors note that part of the PoseBusters V2 comparison with co-folding models is confounded by training-set overlap, which is why they highlight the temporally held-out subset.

What is Tamarind Bio?

Tamarind Bio is a no-code bioinformatics platform built to give life scientists and researchers access to powerful computational tools. Many cutting-edge machine learning models are hard to deploy and use. Tamarind provides an intuitive, web-based environment that removes the complexity of high-performance computing, software dependencies and command-line interfaces.

The platform is designed for biologists, chemists and other researchers who may not have a background in programming or cloud infrastructure but want to run models on their own data. Key features include:

  • A user-friendly graphical interface for setting up and launching experiments

  • A robust API for integration into existing research pipelines

  • An automated system for managing and scaling computational resources

Tamarind treats information and data security as a top priority, as detailed in its Trust Center and Terms of Service.

Accelerating Drug Discovery with MATCHA on Tamarind Bio

  • Structure-based drug design: predict binding poses of candidate ligands in a target protein.

  • Virtual screening: dock large compound libraries where co-folding models would be too slow.

  • Pose quality control: keep poses that are both accurate and physically valid.

  • Pocket-aware docking: supply the binding site when it is known to focus sampling.

How to Use MATCHA on Tamarind Bio

  1. Log in and open the tool: sign in at tamarind.bio and select MATCHA.

  2. Provide the protein: upload the target structure, for example a PDB file.

  3. Provide the ligand: enter a SMILES string or upload a ligand file, as the tool accepts.

  4. Choose the mode: blind docking is the default; pocket-aware docking needs the binding site (if exposed). Choose the full pipeline or the faster LITE variant if offered.

  5. Run the job: submit and wait for sampling, GNINA minimization and filtering to finish.

  6. Download and inspect poses: review the selected pose and its GNINA affinity.

  7. Use downstream: take top poses into rescoring, interaction analysis or further design.

Parameters on Tamarind may differ from the repository, so check the tool page for what is exposed.

Things to Keep in Mind

  • The protein is treated as a rigid body and the ligand is semi-flexible, so large protein conformational changes are not modeled.

  • Poses are ranked by GNINA affinity after minimization, which is a computational score and not a measured binding affinity.

  • Co-folding methods report slightly higher RMSD-only success on some sets, while MATCHA leads on physically valid poses.

  • Results are predictions and need experimental validation.

Source: Frolova, Daulbaev, Sevriugov, Nikolenko, Ivankov, Oseledets and Pak, "MATCHA: Multi-Stage Riemannian Flow Matching for Accurate and Physically Valid Molecular Docking," arXiv preprint 2510.14586 (v3, February 2026). Code and weights: github.com/LigandPro/Matcha.

Supporting 10,000+ scientists around the world,

from leading biotechs, and global biopharma