Predicting molecular properties of molecular variants using residue-specific molecular structural features
Inventors
Shaver, Jeremy Martin • Ketchem, Randal Robert
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
A system for generating a model for predicting a molecular property of a variant of a molecule is provided. For each of a plurality of variants of the molecule, the system for each structural feature, aggregates the values for the structural features of the residues of the molecule that were modified to form the variant to form a feature vector for the variant. The system assigns the value for the molecular property of the variant to the feature vector wherein the feature vector and the assigned value form training data. The system then generates the model for predicting a value for the molecular property using the training data for the plurality of variants.
Core Innovation
The invention describes performing, by a computing system, a training process for a model that predicts values for a molecular property for new variant proteins that correspond to a parent protein. The training process determines sequences of variant proteins having at least one modified residue at a position different from the initial residue at the corresponding position of the parent protein, determines first values for a plurality of structural features with respect to modified residues, and generates structural matrices that indicate a portion of the first values for one or more modified residues of each variant protein.
For each structural matrix, the computing system performs one or more statistical operations on the portion of the first values to produce one or more second values. The system generates, for each variant protein, a structural feature summary matrix that includes the one or more second values, producing a plurality of structural feature summary matrices. Based on the plurality of structural feature summary matrices, the system determines a subset of the plurality of structural features and assigns a value of the molecular property to the modified structural matrix for each variant protein, thereby producing training data that includes the structural feature summary matrices and corresponding molecular property values.
The invention further describes generating, based on the training data, a model that includes one or more parameters corresponding to the subset of the plurality of structural features, and applying the model to new variants. The computing system accesses new variant information indicating one or more additional modified residues at corresponding positions relative to the parent protein, generates an additional structural feature summary matrix for the new variant protein indicating additional values for the subset of structural features, and applies the model to determine an additional value of the molecular property for the new variant protein.
Claims Coverage
The partial content provides three independent claims: a method for training and applying a molecular property prediction model for new variant proteins, a computing system configured to perform the same training and prediction process, and a method that builds and modifies structural matrices for variant molecules and trains and applies a model using those modified matrices. Across the independent claims, the core coverage centers on computing structural-feature values for modified residues, using structural matrices and statistical operations to create structural feature summary matrices, selecting a subset of structural features, and applying a trained model to determine molecular property values for new variants.
Training and predicting molecular property values for new variant proteins with structural feature summary matrices
Performing a training process to predict values for a molecular property for new variant proteins that correspond to a parent protein, including determining sequences of variant proteins with modified residues, determining first values for a plurality of structural features with respect to modified residues, generating structural matrices indicating portions of the first values, performing statistical operations on the portions to produce second values, generating structural feature summary matrices including the second values, determining a subset of the structural features based on the summary matrices, assigning molecular property values to modified structural matrices, producing training data with the summary matrices and the molecular property values, generating a model based on the training data with parameters corresponding to the subset, accessing new variant information indicating additional modified residues, generating an additional structural feature summary matrix indicating additional values for the subset of structural features, and applying the model to determine an additional value of the molecular property for the new variant protein.
Computing system configured for training and applying the structural-feature-based molecular property model
Performing, by a computing system, a training process for a model to predict values for a molecular property for new variant proteins that correspond to a parent protein by determining sequences of variant proteins with modified residues, determining first values for a plurality of structural features with respect to modified residues, generating structural matrices indicating portions of the first values for modified residues, performing statistical operations on portions of the first values to produce second values, generating structural feature summary matrices including the second values, determining a subset of the structural features based on the structural feature summary matrices, assigning values of the molecular property to the modified structural matrices for the variant proteins, producing training data including the structural feature summary matrices and the molecular property values, generating a model based on the training data with parameters corresponding to the subset, accessing input for new variant information indicating additional modified residues, generating an additional structural feature summary matrix for the new variant protein indicating additional values for the subset of structural features, and applying the model to determine an additional value of the molecular property for the new variant protein.
Building structural matrices and modified structural matrices to train and apply the molecular property prediction model
Generating a structural matrix for a variant molecule having modified residues different from initial residues at corresponding positions of a parent molecule, the structural matrix indicating respective first values for individual structural features for individual residues of the variant molecule, modifying the structural matrix to generate a modified structural matrix indicating a subset of the first values that corresponds to the modified individual residues, assigning a value of a molecular property for the variant molecule to the subset of the first values included in the modified structural matrix, generating a model based on the modified structural matrix and the value for the molecular property to predict values for the molecular property for new variant molecules that correspond to the parent molecule, generating a second modified structural matrix for a new variant molecule indicating second values for the individual structural features for individual residues different from individual residues of the parent molecule at corresponding positions, and applying the model to the second modified structural matrix to determine an additional value of the molecular property for the new variant molecule.
Across the independent claims, the inventive concept is implemented through structural feature representations of modified residues, structural matrices and structural feature summary matrices produced by statistical operations, selection of a subset of structural features whose values parameterize the model, and applying the trained model to new variants using additional structural feature summary matrices or modified structural matrices to determine molecular property values.
Stated Advantages
Not explicitly described in patent.
Documented Applications
Not explicitly described in patent.
Interested in licensing this patent?