Use PocketFlow Online
Commercially Available PocketFlow No-Code Web Server
PocketFlow: Structure-Based Generative AI for Drug-Like Molecules Inside Protein Pockets
PocketFlow is a structure-based deep generative model that builds drug-like small molecules directly inside a protein binding pocket, using an autoregressive flow architecture with chemical knowledge built into generation. It was developed by Yuanyuan Jiang, Guo Zhang, Jing You, Hailin Zhang, Rui Yao, Huanzhang Xie, Ziyi Xia, Mengzhe Dai, Yunjie Wu and Shengyong Yang at Sichuan University (West China Hospital and West China School of Pharmacy) and Minjiang University. The code is available on GitHub at Saoge123/PocketFlow.
Finding an active seed compound for a new target is the first and hardest step of a drug discovery project. Conventional high-throughput screening is limited by the structural diversity of existing compound libraries, and ligand-based generative models cannot help when a target has few or no known ligands because they ignore the protein structure entirely.
Earlier structure-based generative models such as LiGAN, GraphBP and Pocket2Mol address part of this, but the paper points to recurring problems: chemically invalid molecules, unusual or unstable rings, poor drug-likeness, unrealistic bond lengths and angles from discretised 3D space, and no wet-lab confirmation that generated molecules are active. PocketFlow was designed to tackle these problems together.
Property | Detail |
|---|---|
Task | Generating small molecules inside a given protein pocket (structure-based de novo design) |
Architecture | Autoregressive normalizing-flow model built on a vector-based equivariant graph network called GDBP |
Input | A protein structure with a defined active pocket |
Training data | About 8 million ZINC molecules for pretraining, then about 150,000 CrossDocked2020 complexes for fine-tuning |
Chemical validity | 100% of generated molecules chemically valid in the paper's benchmark |
Experimental validation | Active inhibitors found for HAT1 and YTHDC1 in the wet lab |
How PocketFlow Generates Molecules in a Protein Pocket
PocketFlow grows a molecule one atom at a time inside the pocket. Each step is split into substeps handled by five modules:
Context Encoder: an equivariant graph attention network encodes the pocket and the molecular fragment built so far.
Focal Net: picks a focal atom that anchors a local coordinate system; atoms with saturated valence are excluded.
Atom Flow: samples the type of the next atom from a learned distribution.
Position Predictor: places the new atom at a continuous 3D position, restricted to within 2 angstroms of the focal atom, so there is no grid discretisation.
Bond Flow: predicts covalent bonds to nearby atoms (within 4 angstroms) using triangular attention, with valence checks and alert-structure filters such as O-O bonds and three-membered rings.
Chemical Knowledge and Transfer Learning
Chemical rules are applied both during and after generation to limit exposure bias, which otherwise accumulates in autoregressive models, particularly for drug-sized molecules above 300 Da. The model also models the covalent bonds of binding-site side chains, and uses transfer learning: pretraining on ligand-only molecules, then fine-tuning on protein-ligand complexes.
PocketFlow Benchmark Results: Drug-Likeness, Synthesizability and Geometry
The authors generated 10,000 molecules for each of 10 test targets and compared PocketFlow with LiGAN, Pocket2Mol and GraphBP, and with real ligands from CrossDocked2020. Averages over the targets:
Metric | CrossDocked2020 | LiGAN | Pocket2Mol | GraphBP | PocketFlow |
|---|---|---|---|---|---|
QED (higher is more drug-like) | 0.531 | 0.480 | 0.397 | 0.415 | 0.507 |
SA score (lower is easier to synthesize) | 3.246 | 4.809 | 4.837 | 5.814 | 2.927 |
Diversity | n/a | 0.898 | 0.855 | 0.897 | 0.877 |
Validity (%) | n/a | 81.4 | 64.8 | 99.6 | 100.0 |
Mean ligand efficiency | n/a | 0.662 | 0.631 | -31.014 (abnormal) | 0.924 |
According to the paper, PocketFlow's bond length and bond angle distributions were closer to real molecules than the baselines (average bond-angle KL divergence 0.228 versus 0.302 for Pocket2Mol), it produced fewer uncommon rings, and its molecules sat mainly inside the pocket. Ligand efficiency was the highest on average and best for 8 of the 10 targets.
Wet-Lab Validation on HAT1 and YTHDC1
The authors applied PocketFlow to two new targets, generating 100,000 molecules each, filtering them by molecular weight, QED and SA score, and synthesizing a few simple candidates.
Target | PDB structure used | Compounds tested | Best result |
|---|---|---|---|
HAT1 (histone acetyltransferase 1) | 6vo5, catalytic site | 2 (H1, H9) | H9, IC50 72.36 +/- 8.03 uM |
YTHDC1 (nuclear m6A reader) | 4r3i, substrate site | 3 (Y1, Y3, Y5) | Y3, IC50 32.6 +/- 2.72 uM; Y5 also active |
The generated binding poses closely matched poses predicted by Glide docking, and the authors describe the hits as low molecular weight, high ligand-efficiency seed compounds.
What is Tamarind Bio?
Tamarind Bio is a no-code bioinformatics platform built to give life scientists and researchers access to powerful computational tools. Many cutting-edge machine learning models are hard to deploy and use. Tamarind provides an intuitive, web-based environment that removes the complexity of high-performance computing, software dependencies and command-line interfaces.
The platform is designed for biologists, chemists and other researchers who may not have a background in programming or cloud infrastructure but want to run models on their own data. Key features include:
A user-friendly graphical interface for setting up and launching experiments
A robust API for integration into existing research pipelines
An automated system for managing and scaling computational resources
Tamarind treats information and data security as a top priority, as detailed in its Trust Center and Terms of Service.
Accelerating Hit Discovery with PocketFlow on Tamarind Bio
Seed compounds for new targets: generate starting molecules for proteins with a known structure but few or no known ligands.
Scaffold ideas beyond screening libraries: explore novel chemotypes that are not limited to existing compound collections.
Prioritisation by drug-likeness: filter generated molecules by QED, synthetic accessibility, molecular weight and ligand efficiency, as the authors did for HAT1 and YTHDC1.
No local setup: run generation without managing GPUs or installing the deep learning stack.
How to Use PocketFlow on Tamarind Bio
Log in and open the tool: sign in at tamarind.bio and select PocketFlow.
Provide the protein structure: upload a PDB file of the target, ideally a crystal structure with a bound ligand or cofactor that marks the pocket.
Define the pocket: specify the active site (for example by a reference ligand or residues), if exposed.
Set generation options: choose how many molecules to generate and any size limits, if exposed. The paper generated 100,000 molecules per target for its case studies.
Run the job: submit and wait for generation to finish.
Download and filter: download the generated molecules and filter by molecular weight, QED and SA score, then rank by docking or ligand efficiency.
Plan follow-up: choose simple, easily synthesized candidates for docking, synthesis and assay, as in the paper.
Parameters shown on Tamarind may differ from the original code, so refer to the tool page for what is exposed.
Things to Keep in Mind
Generated molecules are computational proposals; activity must be confirmed experimentally. The wet-lab hits were micromolar seed compounds, not optimised leads.
Only a handful of compounds were synthesized (2 for HAT1, 3 for YTHDC1), so the validation is a small proof of concept.
Training and generation were limited to molecules containing C, N, O, F, P, S, Cl, Br and I, and fine-tuning complexes had at most 35 heavy atoms.
The results depend on a good pocket definition and on filtering thresholds chosen for each target.
Source: Jiang Y, Zhang G, You J, Zhang H, Yao R, Xie H, Xia Z, Dai M, Wu Y, Yang S. "PocketFlow: an autoregressive flow model incorporated with chemical knowledge for generating drug-like molecules inside protein pockets." Research Square preprint, 2023, DOI 10.21203/rs.3.rs-3077992/v1; the preprint notes publication in Nature Machine Intelligence (2024), DOI 10.1038/s42256-024-00808-8. Code: https://github.com/Saoge123/PocketFlow