Systems and methods for identifying nucleotide sequence matches
Inventors
Blattner, Frederick R. • Baldwin, Schuyler F. • Durfee, Timothy J. • Nash, Daniel A. • Dullea, Kenneth C. • Nelson, Richard D.
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
Systems and methods automatically identify a set of read sequences in one or more larger nucleotide sequences within a set of comparing sequences as a template. The sequences of each set are divided into smaller mer sequences and sorted to arrange the mer sequences in order, and the sets of mers originating from the read sequence set and the comparing sequence set are compared pairwise to determine matching regions between the sequences of the read sequence set and the sequences of the comparing set. The sorting of the sequence sets prior to the pairwise comparison reduces the amount of volatile memory required to assemble the read sequence set and also reduces the overall time to identify matches of the read sequence set in one or more larger nucleotide sequence databases.
Core Innovation
The invention relates to a layout assembly system for identifying nucleotide sequence matches between a read sequence set and a comparing sequence set. The read sequence set comprises a plurality of read entries, each read entry comprising a read oligonucleotide sequence and a read sequence index assigned to identify the corresponding read oligonucleotide sequence. The comparing sequence set comprises a plurality of comparing entries, each comparing entry comprising a comparing oligonucleotide sequence and a comparing sequence index assigned to identify the corresponding comparing oligonucleotide sequence, where the comparing oligonucleotide sequences collectively define a known completely predetermined comparing oligonucleotide sequence in a database of known sequences.
A program stored in non-transitory memory and executable by a processor divides each read oligonucleotide sequence into one or more read mers and assigns each read mer the read sequence index and a read position index identifying a number of nucleotides from a location within the read oligonucleotide sequence. The program generates a read mer table comprising read mer table entries that include the read mer, the read sequence index, and the read position index, and the read mer table entries are sorted by ascending or descending order of read mer sequence. Similarly, the program divides each comparing oligonucleotide sequence into comparing mers, assigns each comparing mer a comparing sequence index and a comparing position index, and generates a comparing mer table that is sorted by ascending or descending order of comparing mer sequence.
The program determines, for each read mer, each match between the respective read mer and a comparing mer in the comparing mer table by comparing the sequence of the read mer to the sequence of a comparing mer in the comparing mer table that has not been previously been compared to the read mer, thereby identifying each match between comparing mers and read mers. The program then orders the read oligonucleotide sequences by sorting the combination of the matches and at least one of the group consisting of sequence index, orientation index, frameshift, and position index of comparing mers, to form an ordered plurality of read oligonucleotide sequences that collectively identify the nucleotide sequence matches in the comparing sequence set.
Claims Coverage
Independent claims are directed to a layout assembly system, a non-transitory computer readable storage medium, and a computer system, each performing a nucleotide sequence match workflow using mer division, indexing and positioning, sorted mer tables, mer-to-mer matching, and ordering based on index, orientation, frameshift, or position information. The claims cover five inventive features shared across these independent claims.
Mer division with read sequence indexing and read position indexing
Divide each respective read oligonucleotide sequence in the read sequence set into one or more read mers, assign a read sequence index to each read mer corresponding to a read oligonucleotide sequence from which the read mer was divided, and assign a read position index to each read mer identifying a number of nucleotides from a location within the read oligonucleotide sequence from which the read mer was divided.
Sorted read mer table generation and ordering by mer sequence
Generate a read mer table comprising read mer table entries stored in non-transitory memory, where each read mer table entry comprises the read mer, the read sequence index, and the read position index for the corresponding read mer, and where the read mer table entries are sorted by ascending or descending order of read mer sequence.
Mer division with comparing sequence indexing and comparing position indexing
Divide each respective comparing oligonucleotide sequence in the comparing sequence set into one or more comparing mers, assign to each comparing mer a comparing oligonucleotide sequence index corresponding to the comparing sequence from which the respective comparing mer was divided, and assign a comparing position index to each comparing mer identifying a number of nucleotides from a location within the comparing oligonucleotide sequence from which the comparing mer was divided.
Sorted comparing mer table generation and matching against not-yet-compared mers
Generate a comparing mer table comprising comparing mer table entries stored in non-transitory memory, where each comparing mer table entry comprises the comparing mer, the comparing sequence index, and the comparing position index for the corresponding comparing mer, where the comparing mer table entries are sorted by ascending or descending order of comparing mer sequence, and determine for each read mer each match between the respective read mer and a comparing mer in the comparing mer table by comparing the sequence of the read mer to the sequence of a comparing mer in the comparing mer table that has not been previously been compared to a read mer in the read mer table.
Ordering read oligonucleotide sequences using match relationships and comparing indices
Order respective read oligonucleotide sequences in the plurality of read oligonucleotide sequences by sorting the combination of the matches between comparing mers and read mers in the comparing mer table and the read mer table, and at least one of the group consisting of sequence index, orientation index, frameshift and position index of comparing mers in the corresponding comparing oligonucleotide sequences, to form an ordered plurality of read oligonucleotide sequences that collectively identify the nucleotide sequence matches in the comparing sequence set.
The independent claims collectively cover mer-based indexing and table construction for both reads and comparing sequences, sorting mer tables by mer sequence, performing comparisons that avoid re-comparing against mers already compared to a read mer, and producing an ordered set of read oligonucleotide sequences based on match results together with at least one of sequence index, orientation index, frameshift, or position index.
Stated Advantages
Not explicitly described in patent.
Documented Applications
Not explicitly described in patent.
Interested in licensing this patent?