System and method for de novo drug discovery
Inventors
Pabrinkis, Aurimas • Bucher, Alwin • Kamuntavi{hacek over (c)}ius, Gintautas • Prat, Alvaro • Bastas, Orestis • Jo{hacek over (c)}ys, {hacek over (Z)}ygimantas • Tal, Roy • Knuff, Charles Dazler
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
A system and method for de novo drug discovery using machine learning algorithms. In a preferred embodiment, de novo drug discovery is performed via data enrichment and interpolation/perturbation of molecule models within the latent space, wherein molecules with certain characteristics can be generated and tested in relation to one or more targeted receptors. Filtering methods may be used to determine active novel molecules by filtering out non-active molecules and contain activity predictors to better navigate the molecule-receptor domain. The system may comprise neural networks trained to reconstruct known ligand-receptors pairs and from the reconstruction model interpolate and perturb the model such that novel and unique molecules are discovered. A second preferred embodiment trains a variational autoencoder coupled with a bioactivity model to predict molecules exhibiting a range of desired properties.
Core Innovation
The invention is directed to a system and a method for de novo drug discovery implemented on a computer system having a memory and a processor. The system receives molecule training data comprising one or more representations of one or more chemical formulas and enriches the molecule training data by querying data sources to find similar molecules or molecules with similar bioactivity. The enriched molecule training data is used to train an encoder of a neural network that determines where each molecule lies in a latent space.
A subspace of the latent space comprising a candidate set of latent examples is determined, and points within the subspace are sampled and subjected to interpolations, perturbations, or both to expand the candidate set of latent examples. The expanded candidate set of latent examples is decoded to reconstruct a candidate set of chemically valid molecules. In related embodiments, the latent-space operations and reconstruction are integrated with a downstream bioactivity module.
In further embodiments, the system uses a voxel-based representation of molecules and trains a variational autoencoder to determine where each molecule lies in a latent space. The decoder portion of the variational autoencoder is used as a vector input to a bioactivity model, and the bioactivity model operates on voxel-based representations to generate concatenated vectors, output compressed latent space, sample compressed latent space, and reconstruct candidate molecules that match desired molecular properties.
Claims Coverage
The partial content identifies three independent claims (clm-00001, clm-00008, clm-00012). Across these independent claims, the inventive features center on latent-space mapping of enriched molecule training data, subspace sampling with interpolation/perturbation and decoding to chemically valid molecules, and voxel-based variational autoencoder representations coupled to a bioactivity model for desired molecular properties.
Latent-space encoder trained on enriched molecule training data
Receive molecule training data comprising one or more representations of one or more chemical formulas; enrich the molecule training data by querying data sources to find similar molecules or molecules with similar bioactivity; use the enriched molecule training data to train an encoder of a neural network, wherein the encoder determines where each molecule in the enriched molecule training data lies in a latent space.
Latent subspace sampling with interpolation/perturbation and decoding to chemically valid molecules
Determine a subspace of the latent space comprising a candidate set of latent examples; sample points within the subspace and perform interpolations, perturbations, or both on the sample points to expand the candidate set of latent examples; decode the candidate set of latent examples to reconstruct a candidate set of chemically valid molecules.
Voxel-based variational autoencoder with bioactivity model vector interface
Receive a voxel-based representation of one or more molecules; use the voxel-based representation of the one or more molecules to train a variational autoencoder, wherein the variational autoencoder determines where each molecule lies in a latent space; use the decoder portion of the variational autoencoder as a vector input to a bioactivity model, wherein the vector input comprises one or more small molecule vector representations.
Concatenated vector compression and reconstruction to match desired molecular properties
Train a model of voxel-based representations on a range of one or more large molecules that comprise desired molecular properties; generate a concatenated vector comprising the one or more small molecule vector inputs and a vector representation of the one or more large molecules; output the concatenated vector to the latent space, wherein the concatenated vector output compresses the latent space; sample the compressed latent space, wherein the compressed latent space comprises a candidate set of latent examples; reconstruct the candidate set of latent examples to arrive at a candidate set of molecules that match the desired molecular properties.
Method steps for latent-space expansion from enriched molecule training data
Receiving molecule training data comprising one or more representations of one or more chemical formulas; enriching the molecule training data by querying data sources to find similar molecules or molecules with similar bioactivity; using the enriched molecule training data to train an encoder of a neural network, wherein the encoder determines where each molecule in the enriched molecule training data lies in a latent space; determining a subspace of the latent space comprising a candidate set of latent examples; sampling points within the subspace and perform interpolations, perturbations, or both on the sample points to expand the candidate set of latent examples; and decoding the candidate set of latent examples to reconstruct a candidate set of chemically valid molecules.
Across the independent claims in the provided content, the core claim coverage is directed to training a neural-network encoder to map enriched molecule representations into a latent space, selecting and expanding a latent subspace via sampling with interpolations/perturbations, and decoding latent examples into chemically valid molecules, with an additional independent claim using a voxel-based variational autoencoder whose decoder vectors feed a bioactivity model to reconstruct molecules matching desired molecular properties.
Stated Advantages
Not explicitly described in patent.
Documented Applications
Not explicitly described in patent.
Interested in licensing this patent?