Implementing a generative machine learning architecture to produce training data for a classification model

Inventors

Shaver, Jeremy MartinAmimeur, TileliKetchem, Randal RobertSmith, Joshua

Assignees

Just Evotec Biologics Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-12080380-B2

Patent

Publication Date

2024-09-03

Expiration Date


Abstract

Amino acid sequences of proteins can be produced using one or more generative machine learning architectures. The amino acid sequences produced by the one or more generative machine learning architectures can be used to train a classification model architecture. The classification model architecture can classify amino acid sequences according to a number of classifications. Individual classifications of the number of classifications can correspond to at least one of a structural feature of proteins, a range of values of a structural feature of proteins, a biophysical property of proteins, or a range of values of a biophysical property of proteins.

Core Innovation

The invention provides a computing-system framework for generating additional amino acid sequences for proteins. It obtains a first training dataset including first amino acid sequences of first proteins, encodes individual first amino acid sequences as a matrix according to amino acids located at positions of the first proteins, and provides generated input data to a generating component of a generative adversarial network.

A generating component produces generated sequences that correspond to additional amino acid sequences, where the generated sequences are represented as vectors indicating amino acids located at a number of positions. A challenging component computationally analyzes the vectors corresponding to generated sequences and the plurality of matrices corresponding to the encoded versions of the first amino acid sequences by implementing a distance function, producing output indicating differences between the generated sequences and the first amino acid sequences.

The training modifies at least one of parameters, weights, or variables of machine learning models of the generating component until a loss function is minimized to produce a trained version of the generating component. The invention then obtains a second training dataset including second amino acid sequences of second proteins with one or more second features of a second group of features that is different from the first group of features, and performs transfer learning with respect to the trained version of the generating component to produce a modified version that produces second additional amino acid sequences having at least one second feature.

The modified version produces third amino acid sequences of third proteins, and an inferential model is trained using at least a portion of the third amino acid sequences to classify amino acid sequences as having features that correspond to at least a portion of the second group of features. Fourth amino acid sequences of fourth proteins are obtained and one or more classifications for the fourth proteins are determined by the trained version of the inferential model, where the classifications indicate at least one or more structural features of the fourth proteins, and in described refinements biophysical properties, with classification probabilities and threshold-based conditions for classification determination, including mapping to ranges of values.

Claims Coverage

The partial document contains three independent claims: method, system, and computer-readable storage media. The inventive coverage combines matrix encoding and vector-represented generation in a generative adversarial network, adversarial distance-function training with loss minimization, transfer learning to a different group of features to generate additional amino acid sequences, and an inferential-model stage that classifies proteins by structural features and, in related refinements, biophysical properties and value ranges.

Encoding amino acid sequences as matrices and generating vector-represented sequences with a generative adversarial network

Encoding individual first amino acid sequences as a matrix according to amino acids located at positions of individual first proteins to produce a plurality of matrices corresponding to encoded versions of the first amino acid sequences; generating input data using a random number generator or pseudo-random number generator provided to a generating component of a generative adversarial network; producing, by the generating component and based on the input data, generated sequences represented as vectors indicating amino acids located at a number of positions.

Challenging component distance-function analysis and loss minimization to train the generating component

Computationally analyzing, by a challenging component of the generative adversarial network and implementing a distance function, the vectors corresponding to the generated sequences and the plurality of matrices corresponding to the encoded versions of the first amino acid sequences to produce output indicating differences between the generated sequences and the first amino acid sequences; modifying, based on the output, at least one of parameters, weights, or variables of one or more machine learning models of the generating component until a loss function of the generating component is minimized to produce a trained version of the generating component.

Transfer learning to a second group of features for generation of additional amino acid sequences

Obtaining a second training dataset including second amino acid sequences of second proteins with one or more second features of a second group of features different from the first group of features; performing transfer learning with respect to the trained version of the generating component based on the second training dataset, including modifying at least one of parameters, weights, or variables in response to minimizing the loss function with respect to the second training dataset to produce a modified version of the generating component; producing additional amino acid sequences of second or third proteins having at least a portion of the second group of features.

Inferential model training and determining classifications indicating structural features

Performing an additional training process for an inferential model using a third training dataset that includes at least a portion of the third amino acid sequences to produce a trained version of the inferential model to classify amino acid sequences as having features that correspond to at least a portion of the second group of features; obtaining fourth amino acid sequences of fourth proteins; determining, by the trained version of the inferential model, one or more classifications for the fourth proteins indicating at least one or more structural features.

Across the independent claims, the inventive coverage is the combination of matrix encoding and vector-represented generation in a generative adversarial network, adversarial distance-function training with loss minimization, transfer learning to a different group of features to generate additional amino acid sequences, and an inferential-model stage that classifies proteins by structural features based on amino acid sequences.

Stated Advantages

Documented Applications

No documented applications found

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.