System and method for prediction of protein-ligand bioactivity and pose propriety
Inventors
Bucher, Alwin • Pabrinkis, Aurimas • Bastas, Orestis • Demtchenko, Mikhail • YANG, Zeyu • Jamieson, Cooper Stergis • Jo{hacek over (c)}ys, {hacek over (Z)}ygimantas • Tal, Roy • Knuff, Charles Dazler
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
A system and method that predicts whether a given protein-ligand pair is active or inactive and outputs a pose score classifying the propriety of the pose. A 3D bioactivity platform comprising a 3D bioactivity module and data platform scrapes empirical lab-based data that a docking simulator uses to generate a dataset from which a 3D-CNN model is trained. The model then may receive new protein-ligand pairs and determine a classification for the bioactivity and pose propriety of that protein-ligand pair. Furthermore, gradients relating to the binding affinity in the 3D model of the molecule may be used to generate profiles from which new protein targets may be determined.
Core Innovation
The invention is a 3D bioactivity platform/system and 3D bioactivity module that trains a machine learning model to predict both protein-ligand bioactivity and pose propriety for a protein-ligand pair. It receives empirical data related to a protein-ligand pair, generates training data from the empirical data, and trains the machine learning model using the generated training data. After training, it receives data related to a target protein-ligand pair and uses the machine learning model to predict bioactivity and the propriety of the pose, then outputs the predictions related to the target protein-ligand pair.
The training can be based on empirical lab-based ground-truth data and docking-simulator-generated training datasets, with training leveraging pose energy/free-energy estimates. The platform generates interpretability outputs including atom-level/binding-affinity gradient visualizations, including an importance reference system/atom-level binding affinity. It also includes an added Pose Score output node to penalize incorrect docking using RMSD-based labels, with training that distinguishes incorrect docking from correct docking via pose propriety labels.
The disclosed approach describes inputs/voxelization using a voxelized cubic grid centered on a binding site, and a pose weighting concept using Boltzmann distributions for poses. It further describes pose generation/alternatives by using force-field-based options for pose generation, and inference options that combine multiple poses. The platform is configured to operate as an integrated data platform, training with empirical data and simulator-generated data to produce predictions of bioactivity and pose propriety.
Claims Coverage
The document includes two independent claims (a system claim and a method claim) that cover receiving empirical data for a protein-ligand pair, generating training data, training a machine learning model, and predicting both protein-ligand bioactivity and pose propriety for a target protein-ligand pair and outputting those predictions. Dependent claim refinements further specify docking-simulator-generated training data, training data comprising one or more poses, and an importance reference system representing binding affinity per atom used during backpropagation.
Predicting protein-ligand bioactivity and pose propriety for a target protein-ligand pair
receiving data related to a target protein-ligand pair; using the machine learning model to predict the bioactivity of the target protein-ligand pair; using the machine learning model to predict the propriety of the pose of the target protein-ligand pair; and output the predictions related to the target protein-ligand pair
Receiving empirical data and generating training data for model training
receive empirical data related to a protein-ligand pair; generate training data from the empirical data; and train a machine learning model using the generated training data
3D bioactivity module driven end-to-end training and output
a 3D bioactivity module comprising instructions that cause the computing device to receive empirical data related to a protein-ligand pair, generate training data from the empirical data, train a machine learning model using the generated training data, receive data related to a target protein-ligand pair, use the machine learning model to predict the bioactivity of the target protein-ligand pair, use the machine learning model to predict the propriety of the pose of the target protein-ligand pair, and output the predictions related to the target protein-ligand pair
Docking-simulator-generated training data
a docking simulator that generates the training data
Training data comprising one or more poses
training data made up of one or more poses of a protein-ligand pair
Importance reference system representing binding affinity per atom
an importance reference system that represents binding affinity for each atom in the molecule
Using importance reference system information during backpropagation
use the importance reference system information during backpropagation to enhance a machine learning model
Across the independent system and method claims, the core coverage is the end-to-end workflow that receives empirical data for a protein-ligand pair, generates training data, trains a machine learning model, and outputs predictions for a target protein-ligand pair covering both bioactivity and pose propriety. The dependent-claim excerpts emphasize docking-simulator-generated training data, training data comprising one or more poses, and an importance reference system representing binding affinity per atom used during backpropagation to enhance the machine learning model.
Stated Advantages
Not explicitly described in patent.
Documented Applications
Not explicitly described in patent.
Interested in licensing this patent?