Use DeCaf Online
Commercially Available DeCaf No-Code Web Server
DeCaf: Few-Step All-Atom Cofolding with Flow Maps for Fast Protein-Ligand Structure Prediction
DeCaf (Denoiser Cofolding All-atom Flowmap) is a framework for distilling all-atom cofolding diffusion models into flow maps that produce high-quality structures in only a few inference steps. It was developed by Gianluca Scarpellini and colleagues at Genesis Molecular AI, MIT, Carnegie Mellon University, Imperial College London and Mila, and described in an arXiv preprint (2026). Code is released on GitHub.
All-atom generative models such as AlphaFold 3 and its successors predict protein and protein-ligand structures at atomic resolution, but they typically need expensive iterative diffusion rollouts. That makes deployment costly and makes inference-time search, where candidate structures are scored by physical reward functions, even more expensive because each reward query needs a fully denoised structure.
DeCaf asks whether a cofolding model can generate reward-optimized samples using only a small number of neural function evaluations (NFEs).
Property | Detail |
|---|---|
Task | Few-step protein-ligand cofolding by distilling diffusion cofolding models |
Teachers | Boltz-1 (DeCaf-Boltz) and Pearl (DeCaf-Pearl) |
Inference method | DeCaf-Search: reward-guided search using flow-map lookahead |
Benchmarks | Runs N' Poses (702 structures) and PoseBusters (282 structures) |
Code | github.com/genesistherapeutics/decaf |
How DeCaf Compresses Diffusion Cofolding into a Few Steps
Flow-map distillation: a denoiser-based flow map is trained so that a few large steps replace a long diffusion trajectory.
SE(3) alignment: endpoint losses support rigid alignment, which the authors show is critical for accurate training.
Sigma-space formulation: a change of variables lets the model work in the noise schedule of EDM-style architectures, enabling direct distillation from pretrained cofolding models.
Flow-map lookahead: the model can estimate the final clean structure from a noisy sample, giving higher-fidelity reward estimates.
DeCaf-Search: candidates are scored by a physical-validity reward and improved with search methods such as Monte Carlo tree search or clean-space gradient ascent.
Results: DeCaf Benchmarks on Runs N' Poses and PoseBusters
On Runs N' Poses, DeCaf-Search was compared with Boltz-1x, which fails at low step counts under its default settings and was therefore tuned for fairness. Success rate means RMSD under 2 Angstrom and PoseBusters-valid.
Budget | Method | PB-valid (%) | Success rate (%) |
|---|---|---|---|
20 NFEs | Boltz-1x (tuned) | 75.9 | 50.3 |
20 NFEs | DeCaf-Search | 83.5 | 57.4 |
40 NFEs | Boltz-1x (tuned) | 86.6 | 55.4 |
40 NFEs | DeCaf-Search | 91.6 | 61.8 |
800 NFEs | Boltz-1x | 95.9 | 64.9 |
800 NFEs | DeCaf-Search (MCTS) | 95.5 | 65.0 |
DeCaf outperforms tuned Boltz-1x at every low budget tested, and matches the per-target RMSD distribution of full Boltz-1x with 5.3x less compute. On PoseBusters it matches full-budget Boltz-1x with 20x less inference compute. When distilling Pearl, DeCaf-Pearl matches its teacher on best@5 success rate (76.0% vs 77.0%, p = 0.593) with 5x fewer NFEs.
What is Tamarind Bio?
Tamarind Bio is a no-code bioinformatics platform built to give life scientists and researchers access to powerful computational tools. Many cutting-edge machine learning models are hard to deploy and use. Tamarind provides an intuitive, web-based environment that removes the complexity of high-performance computing, software dependencies and command-line interfaces.
The platform is designed for biologists, chemists and other researchers who may not have a background in programming or cloud infrastructure but want to run models on their own data. Key features include:
A user-friendly graphical interface for setting up and launching experiments
A robust API for integration into existing research pipelines
An automated system for managing and scaling computational resources
Tamarind treats information and data security as a top priority, as detailed in its Trust Center and Terms of Service.
Accelerating Discovery with DeCaf on Tamarind Bio
Virtual screening: the authors note that the speedup makes it feasible to cofold entire ligand libraries against a target.
Synthetic data generation: faster cofolding yields many more high-quality protein-ligand complexes per unit of compute for training scoring, affinity and generative models.
Reward-guided pose refinement: use physical-validity rewards to steer toward plausible poses at small budgets.
How to Use DeCaf on Tamarind Bio
Open the tool: log in to tamarind.bio and select DeCaf.
Define the complex: provide the protein sequence and the ligand (for example as SMILES), as the interface requests.
Choose the model variant: if exposed, select the Boltz-based or Pearl-based flow map.
Set the budget: if exposed, choose the number of steps or samples; the paper evaluates budgets from 10 to 800 NFEs.
Run: submit the job.
Download outputs: retrieve the predicted complex structures.
Check quality: inspect pocket contacts and physical validity, and compare multiple samples.
Things to Keep in Mind
The reward signal is inherited from Boltz-1x, so failure modes such as non-planar sp2 bonds reflect that choice.
At full Pearl budget the teacher keeps a small edge on single-pose metrics (best@1 success rate and PB-validity).
The relationship between NFE budget and per-target success is non-monotone, suggesting room for adaptive search.
Evaluation focuses on protein-ligand complexes; extension to nucleic acids and larger assemblies is left as future work.
Source: Scarpellini G, Shprints R, Holderrieth P, Nam J, Murugan P, Gomez-Bombarelli R, Jaakkola T, Al-Shedivat M, Boffi NM, Bose AJ. Few-step Cofolding with All-Atom Flow Maps. arXiv:2606.08375 (2026), https://arxiv.org/abs/2606.08375. Code: https://github.com/genesistherapeutics/decaf.