Systems and methods for SNP analysis and genome sequencing

Inventors

Thomas, SterlingDellinger, Nathan

Assignees

Noblis Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-11308056-B2

Patent

Publication Date

2022-04-19

Expiration Date


Abstract

In some embodiments, techniques for identifying one or more species in an undifferentiated environmental sample comprising a plurality of nucleic acid sequences are provided. One or more indices that represent a plurality of reference nucleic acid sequences may be provided, and data may be received comprising digital representations of respective nucleic acid sequences. The respective nucleic acid sequences may be aligned, if possible, using the indices. A respective alignment ration may be calculated for each one of the reference nucleic acid sequences, based on the number of nucleic acid sequences aligned to the respective reference nucleic acid sequence and the total number of nucleic acid sequences in the received data.

Core Innovation

The invention relates to identifying one or more species in an undifferentiated environmental sample using a computer system that aligns received nucleic acid sequences to reference nucleic acid sequences. Reference subsequences are retrieved from positions of reference nucleic acid sequences associated with generated indices, and a hash is computed to determine a corresponding element of the generated index. Position data is stored in the index element reflecting the first position and indicating the reference nucleic acid sequence associated with the generated index.

The system loads one or more indices representing a plurality of reference nucleic acid sequences, where each index comprises elements corresponding to potential nucleic acid sequence permutations. The one or more indices are limited, based on statistical methods regarding which permutations are most likely to occur, to less than a total possible number of permutations. The system receives data comprising data structures digitally representing respective nucleic acid sequences, and for each nucleic acid sequence attempts to align to each reference nucleic acid sequence using the one or more indices.

For each received nucleic acid sequence, the system determines using the one or more indices whether the nucleic acid sequence is successfully aligned to any one or more reference nucleic acid sequences. For each reference nucleic acid sequence, the system calculates a respective alignment ratio based on the number of nucleic acid sequences determined to be successfully aligned to the respective reference nucleic acid sequence and based on the total number of nucleic acid sequences in the received data. The system generates and stores the plurality of calculated alignment ratios for the received data.

Claims Coverage

The identified independent claims are directed to a system, a non-transitory computer-readable storage medium, and a method, each covering the same core approach with generated hash-based indices, indexed alignment attempts, and per-reference alignment ratio calculation. Across the independent claims, the main inventive features number 8.

Species identification from undifferentiated environmental sample

A system, medium, or method for identifying one or more species in an undifferentiated environmental sample comprising a plurality of nucleic acid sequences corresponding to one or more species.

Generated hash-based indices with stored position data

Generate and store at least one generated index of the one or more indices by receiving reference data, identifying a first reference subsequence retrieved from a first position, computing a hash of the first reference subsequence to determine a first corresponding element of the generated index, and storing in that element position data reflecting the first position and indicating the reference nucleic acid sequence associated with the generated index.

Index size limitation using statistical methods

Load one or more indices representing a plurality of reference nucleic acid sequences where each index comprises elements corresponding to a potential nucleic acid sequence permutation, and limit the one or more indices based on statistical methods regarding which permutations are most likely to occur to less than a total possible number of permutations.

Receive nucleic-acid sequence data structures for alignment

Receive data comprising a plurality of data structures, each digitally representing a respective nucleic acid sequence, and for each nucleic acid sequence attempt to align the nucleic acid sequence to each reference nucleic acid sequence using the one or more indices.

Alignment success determination via indexed comparison

For each nucleic acid sequence in the received data, determine using the one or more indices whether the nucleic acid sequence is successfully aligned to any one or more of the reference nucleic acid sequences.

Alignment ratio per reference nucleic acid sequence

For each reference nucleic acid sequence, calculate a respective alignment ratio based on the number of nucleic acid sequences determined to be successfully aligned to the respective reference nucleic acid sequence and based on the total number of nucleic acid sequences in the received data.

Generate and store alignment ratios

Generate and store the plurality of calculated alignment ratios for the received data.

Processor and memory execution

The system, medium, or method comprises a processor and memory and stores instructions executable by the processor to perform the indexed alignment and alignment-ratio calculation.

Across independent claims, the coverage is centered on a processor/memory-based approach that generates and stores hash-based indices that map reference subsequences from positions to index elements containing position data, limits the index contents to a statistically likely set of permutations, uses the indices to attempt alignments for each received nucleic acid sequence and determine alignment success, and computes and stores a per-reference alignment ratio derived from counts of successfully aligned sequences relative to the total received sequences.

Stated Advantages

Not explicitly described in patent.

Documented Applications

Not explicitly described in patent.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.