Autoencoder with generative adversarial network to generate protein sequences

Inventors

Shaver, Jeremy MartinAmimeur, TileliKetchem, Randal Robert

Assignees

Just Evotec Biologics Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-11948664-B2

Patent

Publication Date

2024-04-02

Expiration Date


Abstract

Amino acid sequences of proteins can be produced using an autoencoder. For example, amino acid sequences of variant proteins can be produced by an autoencoder that is fed an amino acid sequence of a base protein as input. A decoding component of the autoencoder can include at least one or more components of a generative adversarial network.

Core Innovation

The invention addresses generating variant protein amino-acid sequences using an autoencoder that includes an encoding component and a decoding component. The decoding component comprises a trained generating component of a generative adversarial network, so the autoencoder is integrated with a GAN-based sequence generation mechanism for producing variant amino-acid sequences from encoded representations.

The training pipeline includes performing a first training process using a first training dataset containing amino acid sequences of first proteins to produce a trained generating component of a generative adversarial network. The invention then produces a second training dataset containing amino acid sequences of second proteins and performs a second training process to generate a trained version of the autoencoder, where the trained version includes a trained encoding component that generates code data representing one or more amino acid sequences of the second training dataset.

After training, the invention provides base sequence data including a first amino acid sequence of a base protein to the trained version of the autoencoder. The autoencoder generates variant sequence data including a second amino acid sequence of a variant protein based on the code data, where the variant sequence includes an amount of similarity and an amount of difference with respect to the base amino-acid sequence.

Claims Coverage

The provided independent claims are clm-00001 and clm-00012. Across these, the inventive coverage centers on training a GAN generating component on protein amino-acid sequences, incorporating that trained GAN generating component as the decoding component of an autoencoder with a trained encoding component, encoding base protein sequence data to produce code data, and generating variant amino-acid sequences with similarity/difference constraints or with differences at one or more positions.

Training a GAN generating component on protein amino-acid sequences

performing a first training process using a first training dataset to produce a trained generating component of a generative adversarial network, the first training dataset including a first plurality of amino acid sequences of first proteins

Autoencoder with GAN generating component as the decoding component

generating an autoencoder that includes an encoding component and a decoding component, the decoding component comprising the trained generating component of the generative adversarial network

Second training to produce a trained autoencoder encoding component

performing a second training process using the second training dataset to generate a trained version of the autoencoder, the trained version of the autoencoder including a trained version of the encoding component that generates code data, the code data representing one or more amino acid sequences of the second training dataset

Generating variant sequence data with similarity and difference to a base sequence

providing base sequence data to the trained version of the autoencoder, the base sequence data including a first amino acid sequence of a base protein; and generating variant sequence data that includes a second amino acid sequence of a variant protein based on the code data, the second amino acid sequence having an amount of similarity with respect to the first amino acid sequence and an amount of difference with respect to the first amino acid sequence

Variant generation via code-data modification and autoencoder decoding with GAN generation

generating code data by an encoding component of an autoencoder, the code data corresponding to a representation of a first amino acid sequence of a base protein that is provided as input to the encoding component; modifying the code data to produce modified code data; providing the modified code data to a decoding component of the autoencoder, the decoding component including a generating component of a generative adversarial network; and generating, by the decoding component, a second amino acid sequence of a variant protein based on the modified code data, the second amino acid sequence having one or more positions with different amino acids than one or more corresponding positions of the first amino acid sequence

The independent claims cover generating variant protein amino-acid sequences by encoding base protein sequences into code data using a trained encoding component, decoding using a GAN generating component within the autoencoder, and producing variant sequences characterized by similarity/difference or by differing amino acids at one or more positions.

Stated Advantages

Not explicitly described in patent.

Documented Applications

Not explicitly described in patent.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.