Use DeCaf Online

Commercially Available DeCaf No-Code Web Server

DeCaf: Few-Step All-Atom Cofolding with Flow Maps for Fast Protein-Ligand Structure Prediction

DeCaf (Denoiser Cofolding All-atom Flowmap) is a framework for distilling all-atom cofolding diffusion models into flow maps that produce high-quality structures in only a few inference steps. It was developed by Gianluca Scarpellini and colleagues at Genesis Molecular AI, MIT, Carnegie Mellon University, Imperial College London and Mila, and described in an arXiv preprint (2026). Code is released on GitHub.

All-atom generative models such as AlphaFold 3 and its successors predict protein and protein-ligand structures at atomic resolution, but they typically need expensive iterative diffusion rollouts. That makes deployment costly and makes inference-time search, where candidate structures are scored by physical reward functions, even more expensive because each reward query needs a fully denoised structure.

DeCaf asks whether a cofolding model can generate reward-optimized samples using only a small number of neural function evaluations (NFEs).

Property

Detail

Task

Few-step protein-ligand cofolding by distilling diffusion cofolding models

Teachers

Boltz-1 (DeCaf-Boltz) and Pearl (DeCaf-Pearl)

Inference method

DeCaf-Search: reward-guided search using flow-map lookahead

Benchmarks

Runs N' Poses (702 structures) and PoseBusters (282 structures)

Code

github.com/genesistherapeutics/decaf

How DeCaf Compresses Diffusion Cofolding into a Few Steps

  1. Flow-map distillation: a denoiser-based flow map is trained so that a few large steps replace a long diffusion trajectory.

  2. SE(3) alignment: endpoint losses support rigid alignment, which the authors show is critical for accurate training.

  3. Sigma-space formulation: a change of variables lets the model work in the noise schedule of EDM-style architectures, enabling direct distillation from pretrained cofolding models.

  4. Flow-map lookahead: the model can estimate the final clean structure from a noisy sample, giving higher-fidelity reward estimates.

  5. DeCaf-Search: candidates are scored by a physical-validity reward and improved with search methods such as Monte Carlo tree search or clean-space gradient ascent.

Results: DeCaf Benchmarks on Runs N' Poses and PoseBusters

On Runs N' Poses, DeCaf-Search was compared with Boltz-1x, which fails at low step counts under its default settings and was therefore tuned for fairness. Success rate means RMSD under 2 Angstrom and PoseBusters-valid.

Budget

Method

PB-valid (%)

Success rate (%)

20 NFEs

Boltz-1x (tuned)

75.9

50.3

20 NFEs

DeCaf-Search

83.5

57.4

40 NFEs

Boltz-1x (tuned)

86.6

55.4

40 NFEs

DeCaf-Search

91.6

61.8

800 NFEs

Boltz-1x

95.9

64.9

800 NFEs

DeCaf-Search (MCTS)

95.5

65.0

DeCaf outperforms tuned Boltz-1x at every low budget tested, and matches the per-target RMSD distribution of full Boltz-1x with 5.3x less compute. On PoseBusters it matches full-budget Boltz-1x with 20x less inference compute. When distilling Pearl, DeCaf-Pearl matches its teacher on best@5 success rate (76.0% vs 77.0%, p = 0.593) with 5x fewer NFEs.

What is Tamarind Bio?

Tamarind Bio is a no-code bioinformatics platform built to give life scientists and researchers access to powerful computational tools. Many cutting-edge machine learning models are hard to deploy and use. Tamarind provides an intuitive, web-based environment that removes the complexity of high-performance computing, software dependencies and command-line interfaces.

The platform is designed for biologists, chemists and other researchers who may not have a background in programming or cloud infrastructure but want to run models on their own data. Key features include:

  • A user-friendly graphical interface for setting up and launching experiments

  • A robust API for integration into existing research pipelines

  • An automated system for managing and scaling computational resources

Tamarind treats information and data security as a top priority, as detailed in its Trust Center and Terms of Service.

Accelerating Discovery with DeCaf on Tamarind Bio

  • Virtual screening: the authors note that the speedup makes it feasible to cofold entire ligand libraries against a target.

  • Synthetic data generation: faster cofolding yields many more high-quality protein-ligand complexes per unit of compute for training scoring, affinity and generative models.

  • Reward-guided pose refinement: use physical-validity rewards to steer toward plausible poses at small budgets.

How to Use DeCaf on Tamarind Bio

  1. Open the tool: log in to tamarind.bio and select DeCaf.

  2. Define the complex: provide the protein sequence and the ligand (for example as SMILES), as the interface requests.

  3. Choose the model variant: if exposed, select the Boltz-based or Pearl-based flow map.

  4. Set the budget: if exposed, choose the number of steps or samples; the paper evaluates budgets from 10 to 800 NFEs.

  5. Run: submit the job.

  6. Download outputs: retrieve the predicted complex structures.

  7. Check quality: inspect pocket contacts and physical validity, and compare multiple samples.

Things to Keep in Mind

  • The reward signal is inherited from Boltz-1x, so failure modes such as non-planar sp2 bonds reflect that choice.

  • At full Pearl budget the teacher keeps a small edge on single-pose metrics (best@1 success rate and PB-validity).

  • The relationship between NFE budget and per-target success is non-monotone, suggesting room for adaptive search.

  • Evaluation focuses on protein-ligand complexes; extension to nucleic acids and larger assemblies is left as future work.

Source: Scarpellini G, Shprints R, Holderrieth P, Nam J, Murugan P, Gomez-Bombarelli R, Jaakkola T, Al-Shedivat M, Boffi NM, Bose AJ. Few-step Cofolding with All-Atom Flow Maps. arXiv:2606.08375 (2026), https://arxiv.org/abs/2606.08375. Code: https://github.com/genesistherapeutics/decaf.

Supporting 10,000+ scientists around the world,

from leading biotechs, and global biopharma