Methods and systems for producing an expanded training set for machine learning using biological sequences

Inventors

Frey, Brendan John • DELONG, Andrew Thomas • XIONG, Hui Yuan

Assignees

Deep Genomics Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-11769073-B2

Patent

Publication Date

2023-09-26

Expiration Date


Abstract

Methods and systems for expanding a training set of one or more original biological sequences are provided. An original training set is obtained, wherein the original training set comprises one or more original biological sequences. Saliency values corresponding to one or more elements in each of the one or more original biological sequences are obtained. For each of the original biological sequences, one or more modified biological sequences are produced and the one or more modified biological sequences are associated with the original biological sequence. One or more elements are generated in each of the one or more modified biological sequences using one or more elements in the associated original biological sequence and the corresponding saliency values. The one or more modified biological sequences for each of the original biological sequences are added to the original training set to form an expanded training set.

Core Innovation

The disclosure relates to computer-implemented methods and systems for training a supervised machine learning model with an expanded training set using biological sequences. The approach obtains a set of original biological sequences comprising deoxyribonucleic acid (DNA) sequences, ribonucleic acid (RNA) sequences, or protein sequences, and obtains saliency values corresponding to one or more elements, where each saliency value indicates a degree of pertinence of the element to biological function. The saliency values are derived from evolutionary conservation across at least two different species, allele frequency in a human population, DNA accessibility, ChIP-Seq, CLIP-Seq, SELEX, a massively parallel reporter assay, or a mutational study.

The disclosure applies one or more modifications to each original biological sequence to create a set of modified biological sequences, while associating each modified biological sequence with an original biological sequence and keeping the same label. It then generates one or more elements in each modified biological sequence using one or more elements from the associated original biological sequence and the corresponding saliency values, and the probability that an element in the modified sequence is the same as the element in the original sequence is higher for larger corresponding saliency values. The biological function of the modified sequences is maintained relative to the associated original biological sequence.

The method creates an expanded training set comprising the original biological sequences and the modified biological sequences, and trains the supervised machine learning model using the expanded training set. In some embodiments, generator parameters represent probabilities over possible values for modified elements, a null symbol denotes a deleted element, and an explicit probability expression uses a saliency transformation and an indicator operator. The supervised machine learning model is trained to determine a molecular phenotype from a biological sequence using labels of molecular phenotypes associated with the original biological sequences.

Claims Coverage

The document contains two independent claims, a computer-implemented method and a computer-implemented system, each covering the same pipeline of obtaining labeled biological sequences, deriving element-level saliency values, generating saliency-guided modified sequences that keep the same label, creating an expanded training set, and training a supervised machine learning model. Dependent claims add specific saliency derivation sources, generator parameters, a null symbol for deleted elements, explicit probability computations, and training for molecular phenotype determination.

Saliency-guided generation of modified biological sequences while preserving labels

obtaining a set of original biological sequences comprising deoxyribonucleic acid (DNA) sequences, ribonucleic acid (RNA) sequences, or protein sequences; obtaining saliency values corresponding to one or more elements in each of the set of original biological sequences, wherein a saliency value indicates a degree of pertinence of the element to biological function; applying one or more modifications to create a set of modified biological sequences and associating each modified biological sequence with an original biological sequence, wherein each original biological sequence has an associated label and each modified biological sequence is associated with the same label as the associated original biological sequence; generating one or more elements in the modified biological sequences using one or more elements in the associated original biological sequence and the corresponding saliency values; and creating an expanded training set comprising the obtained set of original biological sequences and the set of modified biological sequences.

Higher probability of preserving original elements for larger saliency

generating one or more elements in each of the set of modified biological sequences using one or more elements in the associated original biological sequence and the corresponding saliency values, wherein a probability that an element in each of the set of modified biological sequences is the same as the elements in the associated original biological sequence is higher for larger corresponding saliency values, and wherein a biological function of the set of modified biological sequences is maintained relative to the associated original biological sequence.

Training a supervised machine learning model with the expanded training set

training the supervised machine learning model using the expanded training set.

Saliency values derived from biological functional relevance sources

obtaining saliency values corresponding to one or more elements in each of the set of original biological sequences, wherein the saliency values are derived from one or more of evolutionary conservation across at least two different species, allele frequency in a human population of at least two humans, DNA accessibility, chromatin immunoprecipitation sequencing (ChIP-Seq), cross-linking immunoprecipitation sequencing (CLIP-Seq), systematic evolution of ligands by exponential enrichment (SELEX), a massively parallel reporter assay, and a mutational study.

Generator parameters for probabilistic modified-element generation

generating one or more elements in each of the set of modified biological sequences using one or more elements in the associated original biological sequence and the corresponding saliency values, wherein a probability that an element in each of the set of modified biological sequences is the same as the elements in the associated original biological sequence is higher for larger corresponding saliency values; wherein generator parameters are determined and represent probabilities over possible values for modified elements, and the generator parameters are used to generate at least one element in each modified sequence.

Null symbol denoting deleted element

producing a null symbol that denotes a deleted element in at least one modified sequence.

Explicit probability expression using saliency transformation and indicator operator

determining a probability of generating a value alpha for each element xi in modified biological sequences using an indicator operator I(·), a transformation h(si) applied to the saliency value si, and a non-uniform distribution over possible values under constraints.

Saliency derived from evolutionary conservation and human allele frequency

obtaining saliency values derived from evolutionary conservation across at least two different species and allele frequency in a human population of at least two humans.

Supervised learning for molecular phenotype determination

training the supervised machine learning model to determine a molecular phenotype from a biological sequence using labels of molecular phenotypes associated with the original biological sequences.

The claims cover expanding supervised training data for biological sequences by computing element-level saliency values that indicate biological functional pertinence, generating modified sequences by applying modifications whose element-preservation probability increases with saliency, maintaining biological function relative to the originals, associating modified sequences with the same labels, and training a supervised machine learning model on the combined original and modified sequences. Dependent claims add specific saliency derivation sources, generator-parameter mechanisms, deletion via a null symbol, explicit probability computations with saliency transformations, and an objective of determining molecular phenotype.

Stated Advantages

Not explicitly described in patent.

Documented Applications

Not explicitly described in patent.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.