Systems and methods for generating and training convolutional neural networks using biological sequences and relevance scores derived from structural, biochemical, population and evolutionary data
Inventors
XIONG, Hui Yuan • Frey, Brendan
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
We describe systems and methods for generating and training convolutional neural networks using biological sequences and relevance scores derived from structural, biochemical, population and evolutionary data. The convolutional neural networks take as input biological sequences and additional information and output molecular phenotypes. Biological sequences may include DNA, RNA and protein sequences. Molecular phenotypes may include protein-DNA interactions, protein-RNA interactions, protein-protein interactions, splicing patterns, polyadenylation patterns, and microRNA-RNA interactions, which may be described using numerical, categorical or ordinal attributes. Intermediate layers of the convolutional neural networks are weighted using relevance score sequences, for example, conservation tracks. The resulting molecular phenotype convolutional neural networks may be used in genetic testing, to identify drug targets, to identify patients that respond similarly to a drug, to ascertain health risks, or to connect patients that have similar molecular phenotypes.
Core Innovation
The invention relates to molecular phenotype convolutional neural networks (MPCNNs) in which convolutional layer operations are weighted using relevance score sequences. A biological sequence with a plurality of positions is processed by a first layer and subsequent layers of the MPCNN, including convolutional filters that link convolutional layer input positions to convolutional layer output positions. The last layer represents a molecular phenotype.
Each weighting unit is associated with a relevance score sequence comprising relevance score sequence positions, and each relevance score sequence position is associated with a numerical value that quantifies biological relevance of a corresponding position in the biological sequence. The numerical value is defined with respect to the convolutional filters to which the weighting unit is linked. The associated relevance score sequence is used to weight operations of the associated convolutional filter.
The relevance score sequences are derived from structural, biochemical, population, and evolutionary data tracks aligned to biological sequence positions, such as conservation, accessibility, nucleosome, RNA, and protein structural tracks. Relevance score sequences are also generated by a relevance neural network, optionally trained jointly with the MPCNN or trained separately, using backpropagation with modified gradients. The MPCNN uses an encoder and sequence encoding and includes training via gradient-based optimizers.
Claims Coverage
The partial content includes two independent claims: a system claim and a method claim. Both claims cover the same core concept: weighting convolutional filter operations in an MPCNN using relevance score sequences with numerical values that quantify biological relevance per position relative to convolutional filters. Overall, the claims include at least one key inventive feature in each independent claim, with refinements provided by dependent claims.
Position-quantified relevance score sequence weighting of convolutional filters
The MPCNN includes convolutional layers with one or more convolutional filters linking convolutional layer input positions to produced outputs of the convolutional layer, and one or more weighting units each linked to at least one convolutional filter. Each weighting unit is associated with a relevance score sequence comprising relevance score sequence positions, where each numerical value quantifies biological relevance of a corresponding position in the biological sequence with respect to the at least one convolutional filter. Each weighting unit is configured to use the associated relevance score sequence to weight operations of the associated convolutional filter.
Relevance score sequence based weighting operations for convolutional filters in an MPCNN
The method obtains an MPCNN with at least three layers including a first layer obtaining a biological sequence with a plurality of positions and a last layer representing a molecular phenotype, with one or more of the layers configured as convolutional layers having one or more convolutional filters linking received inputs to produced outputs. The method obtains one or more relevance score sequences with numerical values quantifying biological relevance of corresponding positions in the biological sequence with respect to the convolutional filters. The method applies one or more weighting operations in which each weighting operation uses an associated relevance score sequence to weight operations of an associated convolutional filter.
Across both independent claims, the inventive coverage centers on using relevance score sequences that assign numerical biological relevance to biological sequence positions relative to convolutional filters, and applying those relevance scores as weighting operations within an MPCNN to produce molecular phenotype outputs.
Stated Advantages
Not explicitly described in patent.
Documented Applications
genetic testing
drug target/response identification
health risk assessment
patient stratification
Interested in licensing this patent?