Predicting molecular properties of molecular variants using residue-specific molecular structural features

Inventors

Shaver, Jeremy MartinKetchem, Randal Robert

Assignees

Just Evotec Biologics Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-11804283-B2

Patent

Publication Date

2023-10-31

Expiration Date


Abstract

A system for generating a model for predicting a molecular property of a variant of a molecule is provided. For each of a plurality of variants of the molecule, the system for each structural feature, aggregates the values for the structural features of the residues of the molecule that were modified to form the variant to form a feature vector for the variant. The system assigns the value for the molecular property of the variant to the feature vector wherein the feature vector and the assigned value form training data. The system then generates the model for predicting a value for the molecular property using the training data for the plurality of variants.

Core Innovation

The invention describes performing, by a computing system, a training process for a model that predicts values for a molecular property for new variant proteins that correspond to a parent protein. The training process determines sequences of variant proteins having at least one modified residue at a position different from the initial residue at the corresponding position of the parent protein, determines first values for a plurality of structural features with respect to modified residues, and generates structural matrices that indicate a portion of the first values for one or more modified residues of each variant protein.

For each structural matrix, the computing system performs one or more statistical operations on the portion of the first values to produce one or more second values. The system generates, for each variant protein, a structural feature summary matrix that includes the one or more second values, producing a plurality of structural feature summary matrices. Based on the plurality of structural feature summary matrices, the system determines a subset of the plurality of structural features and assigns a value of the molecular property to the modified structural matrix for each variant protein, thereby producing training data that includes the structural feature summary matrices and corresponding molecular property values.

The invention further describes generating, based on the training data, a model that includes one or more parameters corresponding to the subset of the plurality of structural features, and applying the model to new variants. The computing system accesses new variant information indicating one or more additional modified residues at corresponding positions relative to the parent protein, generates an additional structural feature summary matrix for the new variant protein indicating additional values for the subset of structural features, and applies the model to determine an additional value of the molecular property for the new variant protein.

Claims Coverage

The partial content provides three independent claims: a method for training and applying a molecular property prediction model for new variant proteins, a computing system configured to perform the same training and prediction process, and a method that builds and modifies structural matrices for variant molecules and trains and applies a model using those modified matrices. Across the independent claims, the core coverage centers on computing structural-feature values for modified residues, using structural matrices and statistical operations to create structural feature summary matrices, selecting a subset of structural features, and applying a trained model to determine molecular property values for new variants.

Training and predicting molecular property values for new variant proteins with structural feature summary matrices

Performing a training process to predict values for a molecular property for new variant proteins that correspond to a parent protein, including determining sequences of variant proteins with modified residues, determining first values for a plurality of structural features with respect to modified residues, generating structural matrices indicating portions of the first values, performing statistical operations on the portions to produce second values, generating structural feature summary matrices including the second values, determining a subset of the structural features based on the summary matrices, assigning molecular property values to modified structural matrices, producing training data with the summary matrices and the molecular property values, generating a model based on the training data with parameters corresponding to the subset, accessing new variant information indicating additional modified residues, generating an additional structural feature summary matrix indicating additional values for the subset of structural features, and applying the model to determine an additional value of the molecular property for the new variant protein.

Computing system configured for training and applying the structural-feature-based molecular property model

Performing, by a computing system, a training process for a model to predict values for a molecular property for new variant proteins that correspond to a parent protein by determining sequences of variant proteins with modified residues, determining first values for a plurality of structural features with respect to modified residues, generating structural matrices indicating portions of the first values for modified residues, performing statistical operations on portions of the first values to produce second values, generating structural feature summary matrices including the second values, determining a subset of the structural features based on the structural feature summary matrices, assigning values of the molecular property to the modified structural matrices for the variant proteins, producing training data including the structural feature summary matrices and the molecular property values, generating a model based on the training data with parameters corresponding to the subset, accessing input for new variant information indicating additional modified residues, generating an additional structural feature summary matrix for the new variant protein indicating additional values for the subset of structural features, and applying the model to determine an additional value of the molecular property for the new variant protein.

Building structural matrices and modified structural matrices to train and apply the molecular property prediction model

Generating a structural matrix for a variant molecule having modified residues different from initial residues at corresponding positions of a parent molecule, the structural matrix indicating respective first values for individual structural features for individual residues of the variant molecule, modifying the structural matrix to generate a modified structural matrix indicating a subset of the first values that corresponds to the modified individual residues, assigning a value of a molecular property for the variant molecule to the subset of the first values included in the modified structural matrix, generating a model based on the modified structural matrix and the value for the molecular property to predict values for the molecular property for new variant molecules that correspond to the parent molecule, generating a second modified structural matrix for a new variant molecule indicating second values for the individual structural features for individual residues different from individual residues of the parent molecule at corresponding positions, and applying the model to the second modified structural matrix to determine an additional value of the molecular property for the new variant molecule.

Across the independent claims, the inventive concept is implemented through structural feature representations of modified residues, structural matrices and structural feature summary matrices produced by statistical operations, selection of a subset of structural features whose values parameterize the model, and applying the trained model to new variants using additional structural feature summary matrices or modified structural matrices to determine molecular property values.

Stated Advantages

Not explicitly described in patent.

Documented Applications

Not explicitly described in patent.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.