Ordinal position-specific and hash-based efficient comparison of sequencing results

Inventors

Trooskens, Geert • Van Criekinge, Wim Maria R.

Assignees

DocAI Inc • Sharecare AI Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-11551784-B2

Patent

Publication Date

2023-01-10

Expiration Date


Abstract

The technology disclosed generates a reference array of variant data for locations that are shared between read results which are to be compared, and generates hashes over a selected pattern length of positions in the reference array to independently produce non-unique window hashes for base patterns in the read results. It then selects for comparison window hashes that occur less than a ceiling number of times and compares the selected window hashes to identify common window hashes between the read results. It then determines a similarity measure for the read results based on the common window hashes.

Core Innovation

The invention relates to efficiently comparing sequenced outputs by accessing a first sequenced output and a second sequenced output that contain variants occurring at different carriers and at different carrier positions. A reference array is generated for carrier positions that are shared between the first and second sequenced outputs, and a first sequence is generated from the first sequenced output and a second sequence is generated from the second sequenced output based on the reference array. This framework focuses the comparison on carrier positions shared across the two sequenced outputs.

Hashes are generated over a selected pattern length of positions in the reference array to independently produce non-unique window hashes for base patterns in the first and second sequences. For comparison, window hashes are selected based on occurrences that are less than a ceiling number of times. The selected window hashes are compared between the first and second sequences on a starting position basis so that selected window hashes for base patterns having same start positions in the read results are compared.

Common window hashes are identified between the first and second sequences based on the comparing, and a similarity measure between the first and second sequences is determined based on the common window hashes. The approach enables similarity measurements derived from common, occurrence-filtered window hashes while maintaining alignment by starting position basis.

Claims Coverage

The provided material includes three independent claims. Each independent claim centers on a shared carrier-position reference array, non-unique window hashing over a selected pattern length, ceiling-based selection of hashes, starting-position-based comparison, identification of common window hashes, and determination of a similarity measure from those common hashes.

Shared carrier-position reference array for aligning sequenced outputs

Accessing a first sequenced output and a second sequenced output that contain variants occurring at different carriers and at different carrier positions, generating a reference array for carrier positions that are shared between the first and second sequenced outputs, and generating a first sequence from the first sequenced output and a second sequence from the second sequenced output based on the reference array.

Non-unique window hashes over selected pattern length

Generating hashes over a selected pattern length of positions in the reference array to independently produce non-unique window hashes for base patterns in the first and second sequences.

Ceiling-based selection of window hashes

Selecting for comparison window hashes that occur less than a ceiling number of times.

Starting position basis comparison of window hashes

Comparing the selected window hashes between the first and second sequences on a starting position basis such that selected window hashes for base patterns having same start positions in the read results are compared.

Common window hash identification and similarity determination

Identifying common window hashes between the first and second sequences based on the comparing, and determining a similarity measure between the first and second sequences based on the common window hashes.

Computer-implemented method stored on a non-transitory medium

A non-transitory computer readable storage medium impressed with computer program instructions that, when executed on a processor, implement actions including accessing first and second sequenced outputs, generating a shared carrier-position reference array, generating first and second sequences based on the reference array, generating non-unique window hashes over a selected pattern length, selecting window hashes using a ceiling number of times, comparing selected window hashes on a starting position basis, identifying common window hashes, and determining a similarity measure based on the common window hashes.

System for processor-executed sequenced-output comparison

A system including one or more processors coupled to memory, the memory loaded with computer instructions to implement actions including accessing first and second sequenced outputs, generating a shared carrier-position reference array, generating first and second sequences based on the reference array, generating non-unique window hashes over a selected pattern length, selecting window hashes using a ceiling number of times, comparing selected window hashes on a starting position basis, identifying common window hashes, and determining a similarity measure based on the common window hashes.

Across the independent claims, the inventive core is a shared carrier-position reference array that drives generation of first and second sequences, creation of non-unique window hashes over a selected pattern length, selection of hashes whose occurrence is below a ceiling, starting-position-basis comparison, identification of common window hashes, and determination of a similarity measure based on those common window hashes.

Stated Advantages

Efficiently comparing sequenced outputs.

Documented Applications

Not explicitly described in patent.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.