Method and system for rapid searching of genomic data and uses thereof
Inventors
Daly, Kristopher Edward • Stephens, Kurtis Phillip
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
A method, apparatus and system for transforming genomic data into a computer database environment comprising a forward lookup table and a plurality of reverse lookup tables which relate consecutive overlapping reference sequence segments to reference sequences stored in the forward lookup table enables rapid and precise matching of undefined biological sequences with reference sequences.
Core Innovation
The method receives genomic data comprising a plurality of reference sequences and constructs, via a data processing system, a forward lookup table and at least one reverse lookup table. Each reference sequence is transformed into a plurality of segments, and constructing the forward lookup table parses the genomic data to create a reference sequence record for each reference sequence in the genomic data.
Constructing the at least one reverse lookup table transforms each reference sequence record in the forward lookup table into a set of Y−k+1 consecutive overlapping reference sequence segments and stores each transformed reference sequence segment in a record in the at least one reverse lookup table. Each reverse lookup table record contains a reference sequence segment, a pointer indicating which reference sequence record contains the reference sequence segment, and an offset (I) indicating a position at which the reference sequence segment begins in the second string. Each reference sequence record comprises a column indicating a species of origin associated with the reference sequence.
For matching, the method receives a query comprising at least one query sequence having a third string of contiguous characters having a string length equal to L and selects at least a first reverse lookup table that has the largest preselected positive integer k that is less than or equal to L. The method partitions the third string into L/k consecutive non-overlapping query sequence segments and queries each query segment against each set of Y−k+1 consecutive overlapping reference sequence segments in the selected reverse lookup table to obtain matching reference segments.
The method outputs at least one reference sequence that matches the at least one query sequence, including reconstruction of matched regions using species of origin, sorting by offsets, and concatenating segments, with additional reverse lookup tables used for remaining gaps using smaller k values.
Claims Coverage
The provided content includes one explicit independent claim, supported by dependent refinements. The independent claim includes four core inventive features related to lookup-table construction from reference sequences using overlapping k-length segments, species-of-origin-aware storage using pointers and offsets, query-time selection of an appropriate reverse lookup table followed by segment partitioning and matching, and output of matching reference sequences.
Bidirectional lookup-table construction with overlapping k-length reverse segments
Constructing a forward lookup table and at least one reverse lookup table by transforming each reference sequence into records, where the reverse lookup table stores Y−k+1 consecutive overlapping reference sequence segments of first strings having string length k corresponding to contiguous k characters of the second string, with pointers to forward-table records.
Reverse lookup records with offsets and species-of-origin column
Each reverse lookup table record contains a pointer and an offset (I) indicating the segment start position in the reference sequence string, and each reference sequence record comprises a column indicating a species of origin.
Selecting the reverse lookup table based on the largest k ≤ L and partitioning the query
Selecting at least a first reverse lookup table having the largest preselected positive integer k that is less than or equal to the query string length L, and transforming the query sequence into at least a first set of L/k consecutive non-overlapping query sequence segments by partitioning the query string into L/k non-overlapping segments.
Segment-by-segment matching against overlapping reverse segments and outputting matches
Querying each non-overlapping query segment against each set of Y−k+1 consecutive overlapping reference sequence segments in the selected reverse lookup table for matching, and outputting at least one reference sequence that matches the at least one query sequence.
Species-filtered reconstruction by sorting matched segments by offsets and concatenating
Reconstructing matching reference sequence segments for each species of origin by filtering query results by species, sorting segments for each species by the first offset (I1) in ascending order, and concatenating the sorted segments to form a fourth string of contiguous characters.
Multi-stage gap filling using additional reverse lookup tables with smaller k and optional sliding, permutative, or hybrid matching
Selecting at least a second reverse lookup table with a smaller preselected integer k2 than k and using sliding matching, permutative matching, or hybrid matching to fill remaining gaps in a fourth contiguous-character string.
Report output including alignment, species histogram, and insertion, deletion, or mutation polymorphism identification
Outputting a report including matched reference sequences and one or more of: a histogram indicating likely species of origin, an alignment to forward lookup table reference sequences, and identification of one or more insertion, deletion, or mutation polymorphisms.
The coverage centers on constructing forward and reverse lookup tables using overlapping k-length segments with pointers and offsets, selecting a reverse lookup table based on k relative to query length L, partitioning the query into L/k non-overlapping segments for matching, and reconstructing matched regions using species-of-origin filtering and offset-sorted concatenation, with additional gap filling using smaller k and report outputs that can include species histograms, alignments, and insertion, deletion, or mutation polymorphism identification.
Stated Advantages
Substantial runtime improvements are reported versus SSAHA and BLAST.
Documented Applications
BRCA1 SNP detection (Rs16942).
Virus homology and species differentiation for a protein sequence.
Interested in licensing this patent?