This model is a graph Transformer foundation model that jointly learns geometric and chemical representations of protein pockets and ligands from large-scale 3D structural data. MolX integrates over 3 million protein pockets and 5 million molecules, representing both entities as E(3)-equivariant graphs that preserve spatial geometry and chemical context.
-
Prepare the data. To pretrain MolX, users are supposed to download the pretrain datasets including MoltextNet, Pcqm4m-v2, and PDB pocket from .
-
Place files. Place the downloaded pretrained data in the corresponding paths in
examplex/property_prediction/dataset. -
Prepare the environment. Here we export our anaconda environment as the file "env.yaml". You can use the command:
conda env create -f env.yaml conda activate MOLX
to get the same environment. Also, we use one A100 GPU with Ubuntu 18.04 to conduct pretraining.
Besides, we highly recommond to install openbabel (2.3.2) (https://openbabel.org/wiki/Main_Page) and preprocess the mol2 files.
apt install openbabel
-
Run the training script.
sh multiloss.sh
Please run the corresponding shell files in examples/property_prediction to finetune different datasets.
For example, to finetune MolX on Molecule Glue dataset, run the following script.
sh mg.sh