Use CryoLigATE Online
Commercially Available CryoLigATE No-Code Web Server
CryoLigATE: Deep Learning to Enhance Ligand Density in Cryo-EM Protein-Ligand Maps
CryoLigATE (cryo-EM ligand AI-trained enhancement) is a deep learning framework that enhances the density of protein-bound ligands in cryo-EM maps. It was developed by Nandan Haloi, Rebecca J. Howard and Erik Lindahl at Stockholm University, KTH Royal Institute of Technology and collaborating institutions. The source code and pretrained weights are hosted on GitHub.
Cryo-EM has become central to structure-based drug discovery, yet ligand-binding sites are often resolved less well than the surrounding protein. The authors cite a beta-galactosidase example in which the global protein density reached 1.5 Å while the ligand density resolved only to 3.0 to 3.5 Å. Automated model-building tools depend on map quality, so a poorly resolved ligand limits confident atomic interpretation.
Existing AI map post-processing methods such as EMReady and DeepEMhancer were trained largely on protein maps. According to the authors, they are less effective for ligands and lipids, which cover a broader and more irregular chemical space, and treating maps with them can lower model-to-map Q-scores at the binding site.
Property | Detail |
|---|---|
Task | Local enhancement of ligand density around a binding pocket |
Architecture | 3D Swin-Conv UNet (hybrid convolution and transformer) |
Training data | 6,511 protein-ligand pairs from cryo-EM structures, resolution of at least 4 Å |
Inputs | Cryo-EM map plus an aligned PDB model, with ligand residue name and number |
Speed | About a second per complex on a desktop RTX 3070 GPU |
Availability | Code and weights on GitHub; dataset and training code on Zenodo |
How CryoLigATE Improves Ligand Resolvability in Cryo-EM Maps
Automatic pocket extraction: the pipeline uses the coordinates of a preliminary atomic model only to locate the region of interest and crops a localized sub-volume, so no manual map preparation is needed.
Network input: only the localized experimental density and a protein occupancy mask are given to the model, which does not use SMILES strings or chemical identifiers.
Training on simulated targets: maps were resampled to 64 x 64 x 64 voxel volumes (32 Å cubes), and forward-simulated maps from deposited atomic models served as the high-resolution ground truth.
Sharpened output: the model returns an enhanced density map of the binding pocket that can be used for ligand and protein model building.
A Diverse Protein-Ligand Training Set
The dataset spans drug-like molecules (59.3%), lipids (12.2%) and nucleotides (9.7%), and proteins including receptors (22.1%), transporters (10.1%), ion channels (4.6%), translation machinery (13.3%) and enzymes (5.2%). It was split by sequence-identity clustering into 5,195 training, 667 validation and 649 test complexes.
Benchmark Results on 649 Independent Test Complexes
Subset | Result |
|---|---|
All test complexes (n = 649) | No significant Q-score gain for the overall pocket or ligand-only density |
Poor input ligand density, Q-score below 0.5 (n = 81) | Significant ligand Q-score improvement |
Well resolved input, Q-score above 0.5 (n = 558) | No significant reduction in Q-score |
Change in ligand Q-score | Broadly positive, from -0.05 to 0.3 |
The authors conclude that CryoLigATE helps most when the ligand density is poor and does not degrade high-quality maps. In a perturbation test on the GABAA receptor bound to allopregnanolone, it recovered the expected ligand density despite a 90 degree rotation, a 3 Å shift or replacement of the ligand by a single carbon atom in the input model. Enhanced maps showed clearer functional groups and more continuous density.
What is Tamarind Bio?
Tamarind Bio is a no-code bioinformatics platform built to give life scientists and researchers access to powerful computational tools. Many cutting-edge machine learning models are hard to deploy and use. Tamarind provides an intuitive, web-based environment that removes the complexity of high-performance computing, software dependencies and command-line interfaces.
The platform is designed for biologists, chemists and other researchers who may not have a background in programming or cloud infrastructure but want to run models on their own data. Key features include:
A user-friendly graphical interface for setting up and launching experiments
A robust API for integration into existing research pipelines
An automated system for managing and scaling computational resources
Tamarind treats information and data security as a top priority, as detailed in its Trust Center and Terms of Service.
Accelerating Discovery with CryoLigATE on Tamarind Bio
Structure-based drug design: sharpen ligand density to build more confident models of drug-like molecules in cryo-EM structures.
Lipids and steroids: improve interpretation of lipid, steroid and carbohydrate densities in membrane proteins.
Weak ligand sites: recover functional groups and continuity for ligands in maps with poorly resolved pockets.
Model validation: compare enhanced and original maps to check ligand placement and conformation.
How to Use CryoLigATE on Tamarind Bio
Log in and open the tool: sign in at tamarind.bio and select CryoLigATE.
Upload the cryo-EM map: provide the experimental volume in map format, as the command-line tool expects.
Upload an aligned atomic model: provide a PDB file containing the ligand, aligned to the map, which is used only to locate the pocket.
Specify the ligand: enter the ligand residue name and residue number, if exposed.
Run the enhancement: submit the job; the paper reports about a second per complex on a desktop GPU.
Download the enhanced map: retrieve the sharpened pocket density.
Rebuild and refine: open the original and enhanced maps in your modelling software and use the enhanced density for ligand model building, then validate against the original map.
Things to Keep in Mind
CryoLigATE should refine existing, putative ligand density only. In a map region with no density it produced amorphous density, so it should not be used on visibly empty pockets.
Gains were not significant across all test complexes or for surrounding protein residues; benefits appear mainly for poorly resolved ligands.
The ligand is treated as a single static entity, so heterogeneous binding and low-occupancy states are not modelled.
High-resolution cryo-EM protein-ligand training data are limited, and deposited models may contain coordinate errors.
Source: Haloi, N., Howard, R. J. and Lindahl, E., "CryoLigATE: enhancing the resolvability of cryo-EM maps in protein-ligand complexes using deep learning," bioRxiv preprint, August 2026, doi:10.64898/2026.08.04.742718. Code: github.com/nandanhaloi123/CryoLigATE. Data: Zenodo, doi:10.5281/zenodo.20794212.