Bin-specific and hash-based efficient comparison of sequencing results
Inventors
Trooskens, Geert • Van Criekinge, Wim Maria R.
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
The technology disclosed generates a reference array of variant data for locations that are shared between read results which are to be compared, and generates hashes over a selected pattern length of positions in the reference array to independently produce non-unique window hashes for base patterns in the read results. It then selects for comparison window hashes that occur less than a ceiling number of times and compares the selected window hashes to identify common window hashes between the read results. It then determines a similarity measure for the read results based on the common window hashes.
Core Innovation
The invention relates to efficiently comparing sequenced outputs. The sequenced outputs contain variants occurring at different carriers and at different carrier positions and are partitioned into bins. A reference array is generated for those carrier positions that are shared between the first and second sequenced outputs, and a first sequence and a second sequence are generated from the sequenced outputs based on the reference array.
Hashes are generated over a selected pattern length of positions in the reference array to independently produce non-unique window hashes for base patterns in the first and second sequences. Window hashes are selected for comparison based on occurrence frequency, specifically selecting window hashes that occur less than a ceiling number of times. The selected window hashes are compared between the first and second sequences on a bin-by-bin basis, where comparison is restricted to window hashes produced for base patterns within the given bin in each sequence.
Common window hashes are identified for each bin in the first and second sequences based on the comparing. A similarity measure for each bin is determined based on the common window hashes. A distance formula is disclosed to determine similarity using common window hashes and unique window hashes, including derived outputs such as a distance tree, as well as downstream uses including inherited-trait relatedness inference, ethnic ancestry/origination estimation, and comparative genomic browser visualizations.
Claims Coverage
The independent claims include three independent claim types (method, non-transitory computer readable storage medium, and system) covering the same core inventive workflow. Across these independent claims, five main inventive features are consistently present: reference-array generation for shared carrier positions, sequence generation from sequenced outputs, non-unique window hashing over a selected pattern length, frequency-thresholded hash selection, and bin-by-bin comparison to identify common window hashes and compute per-bin similarity measures.
Shared-carrier reference array alignment for partitioned bins
accessing a first sequenced output and a second sequenced output, wherein the first and second sequenced outputs contain variants occurring at different carriers and at different carrier positions and are partitioned into bins; generating a reference array for those carrier positions that are shared between the first and second sequenced outputs;
Reference-array-based sequence generation
based on the reference array, generating a first sequence from the first sequenced output and a second sequence from the second sequenced output;
Non-unique window hashes over a selected pattern length
generating hashes over a selected pattern length of positions in the reference array to independently produce non-unique window hashes for base patterns in the first and second sequences;
Frequency ceiling selection of window hashes
selecting for comparison window hashes that occur less than a ceiling number of times;
Bin-by-bin common window hash identification and bin similarity determination
comparing the selected window hashes between the first and second sequences on a bin-by-bin basis such that a first set of selected window hashes produced for base patterns in a given bin in the first sequenced output are compared only to a second set of selected window hashes produced for base patterns in the given bin in the second sequenced output; identifying common window hashes for each bin in the first and second sequences based on the comparing; and determining a similarity measure for each bin based on the common window hashes.
All independent claims focus on efficiently comparing sequenced outputs by aligning shared carrier positions into a reference array, generating sequences, producing non-unique window hashes over a selected pattern length, selecting hashes under a ceiling occurrence threshold, and performing bin-by-bin comparisons to identify common window hashes and compute per-bin similarity measures.
Stated Advantages
Efficiently comparing sequenced outputs by using reference-array-based hashing and bin-by-bin comparison.
Restricting comparisons to common window hashes within corresponding bins.
Enabling determination of a similarity measure for each bin based on common window hashes.
Documented Applications
Similarity-based distance trees derived from bin-wise similarity measures.
Inherited-trait relatedness inference based on shared-base percentage and similarity measures [procedural detail omitted for safety].
Ethnic ancestry/origination estimation based on similarity measures [procedural detail omitted for safety].
Comparative genomic browser visualizations using the comparison results [procedural detail omitted for safety].
Interested in licensing this patent?