Hash-based efficient comparison of sequencing results
Inventors
Trooskens, Geert • Van Criekinge, Wim Maria R.
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
First and second sequenced outputs are accessed. The sequenced outputs contain variants occurring at different carriers and at different carrier positions. Hashes are generated over a selected pattern length of positions for those carrier positions that are shared between the sequenced outputs to produce window hashes for base patterns in first and second sequences. Each sequence is based on the shared carrier positions and the respective sequenced output. The window hashes are non-unique. Window hashes that occur less than a ceiling number times are selected. The selected window hashes are compared between the sequences on a starting position basis such that selected window hashes for base patterns having same start positions in the sequenced outputs are compared. Common window hashes are identified between the sequences based on the comparing. A similarity measure is determined between the sequences based on the common window hashes.
Core Innovation
The invention provides a computer-implemented method, and a corresponding system and storage medium with instructions, for comparing a first sequenced output and a second sequenced output that contain variants occurring at different carriers and at different carrier positions. Shared carrier positions between the first and second sequenced outputs are used to generate window hashes over a selected pattern length of positions, producing window hashes for base patterns in a first sequence and a second sequence based on the shared carrier positions.
Window hashes are non-unique, and the method selects those window hashes that occur less than a ceiling number of times. The selected window hashes are compared between the first sequence and the second sequence on a starting position basis such that selected window hashes for base patterns having same start positions in the first sequenced output and the second sequenced output are compared.
Common window hashes between the first sequence and the second sequence are identified based on the comparing, and a similarity measure is determined between the first sequence and the second sequence based on the common window hashes. The described implementations support similarity and distance characterization, including carrier-by-carrier analysis, and can be visualized using a distance tree visualization and a genomic browser visualization.
Claims Coverage
Independent claims are provided for a computer-implemented method, a non-transitory computer readable storage medium with program instructions, and a system. Each independent claim centers on hashing non-unique window patterns over shared carrier positions, filtering by ceiling frequency, comparing by matching start positions, identifying common window hashes, and determining a similarity measure based on the common window hashes.
Non-unique window hash comparison on shared carrier positions
Accessing a first sequenced output and a second sequenced output containing variants at different carriers and carrier positions; generating hashes over a selected pattern length of positions for those carrier positions shared between the first and second sequenced outputs to produce window hashes for base patterns in a first sequence and a second sequence based on the shared carrier positions, where the window hashes are non-unique.
Ceiling-frequency filtering of window hashes
Selecting those of the window hashes that occur less than a ceiling number of times.
Starting position basis comparison
Comparing the selected window hashes between the first sequence and the second sequence on a starting position basis such that selected window hashes for base patterns having same start positions in the first sequenced output and the second sequenced output are compared.
Common window hash identification and similarity measurement
Identifying common window hashes between the first sequence and the second sequence based on the comparing; determining a similarity measure between the first sequence and the second sequence based on the common window hashes.
Non-transitory program instructions for window hash similarity
A non-transitory computer readable storage medium impressed with computer program instructions that, when executed on a processor, implement a method comprising accessing first and second sequenced outputs, generating non-unique window hashes over shared carrier positions and a selected pattern length, selecting window hashes occurring less than a ceiling number of times, comparing selected window hashes by matching start positions, identifying common window hashes, and determining a similarity measure based on the common window hashes.
Processor-implemented system for window hash similarity
A system comprising one or more processors coupled to memory, where memory stores computer instructions that implement actions comprising accessing first and second sequenced outputs, generating non-unique window hashes over shared carrier positions for a selected pattern length, selecting window hashes occurring less than a ceiling number of times, comparing selected window hashes on a starting position basis, identifying common window hashes, and determining a similarity measure based on the common window hashes.
Across all independent claims, the core inventive concept is computing similarity between two sequenced outputs by generating non-unique window hashes over shared carrier positions, filtering by ceiling-frequency occurrence, comparing by matching start positions, identifying common window hashes, and determining a similarity measure based on those common window hashes.
Stated Advantages
Not explicitly described in patent.
Documented Applications
Not explicitly described in patent.
Interested in licensing this patent?