Methods for detecting nucleic acid variants
Inventors
ETZIONI, YOAV • Faigler, Simchon • Almogy, Gilad • Pratt, Mark • Oberstrass, Florian
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
Methods for detecting a short genetic variant in a test sample are described herein. In some exemplary methods, the short genetic variant is called using one or more match scores, which are determined using one or more sequencing data sets obtained from a test nucleic acid molecule, wherein the test sequencing data sets are determined by sequencing the test nucleic acid molecule using non-terminating nucleotides provided in separate nucleotide flows according to a flow-cycle order. Also described herein are methods of sequencing a test nucleic acid molecule using two or more different flow-cycle orders and/or extended flow cycle orders having five or more nucleotide flows per flow cycle.
Core Innovation
The invention provides a method for detecting a short genetic variant associated with a disease in a test sample by using sequencing with non-terminating nucleotides provided in separate flow positions according to a flow-cycle order. A target short genetic variant associated with a disease is selected such that a target sequencing data set differs from a reference sequencing data set at four or more consecutive flow positions. Sequencing is performed for the target sequence and the reference sequence to obtain the target sequencing data set and the reference sequencing data set.
The method obtains one or more test sequencing data sets derived from the test sample by sequencing test nucleic acid sequences at least partially overlapping a locus associated with the target short genetic variant. For each test nucleic acid sequence, a respective match score is determined that indicates a likelihood that the test sequencing data set matches the target sequencing data set or indicates a likelihood that the test sequencing data set matches the reference sequencing data set. The target short genetic variant is selected prior to calling the presence or absence of the target short genetic variant in the test sample.
Using the one or more respective match scores, the method calls the presence or absence of the target short genetic variant in the test sample. In related aspects, one or more first test sequencing data sets are obtained using a first flow-cycle order and one or more second test sequencing data sets are obtained by re-sequencing the same test nucleic acid molecules using a different second flow-cycle order, and respective match scores to candidate sequences are determined based on likelihood. Further aspects include sequencing a nucleic acid molecule by hybridizing to a primer, extending using labeled non-terminating nucleotides in a repeated flow-cycle order with five or more separate nucleotide flows, and detecting a signal from an incorporated labeled nucleotide or an absence of a signal as the primer is extended.
Claims Coverage
The provided independent claims are clm-00001, clm-00018, and clm-00019. Across these claims, the core claim coverage includes three main inventive areas: disease-associated detection of a selected target short genetic variant using flow-position match-score likelihoods; two-pass re-sequencing with different flow-cycle orders to support likelihood-based calling; and sequencing-by-synthesis using labeled non-terminating nucleotides with signal/absence detection under a minimum repeated flow-cycle order size.
Selecting a disease-associated target short genetic variant with flow-position difference
Selecting a target short genetic variant associated with a disease such that a target sequencing data set associated with a target sequence comprising the target short genetic variant differs from a reference sequencing data set associated with a reference sequence at four or more consecutive flow positions when obtained by sequencing the target sequence and reference sequence using non-terminating nucleotides provided in separate flow positions according to a flow-cycle order.
Obtaining test sequencing data sets from overlapping loci
Obtaining one or more test sequencing data sets, each test sequencing data set associated with a test nucleic acid sequence, each test nucleic acid sequence at least partially overlapping a locus associated with the target short genetic variant and derived from the test sample, wherein the one or more test sequencing data sets were determined by sequencing the test sample using non-terminating nucleotides provided in separate flow positions according to the flow-cycle order.
Determining likelihood-indicative match scores for target or reference
Determining, for each test nucleic acid sequence, a respective match score indicative of a likelihood that the test sequencing data set matches the target sequencing data set, or a respective match score indicative of a likelihood that the test sequencing data set matches the reference sequencing data set.
Calling presence or absence based on match scores
Calling, using the one or more respective match scores, the presence or absence of the target short genetic variant in the test sample, where the target short genetic variant is selected prior to calling the presence or absence.
Two-pass re-sequencing using different flow-cycle orders
Obtaining one or more first test sequencing data sets determined by sequencing test nucleic acid molecules using non-terminating nucleotides provided in separate flow positions according to a first flow-cycle order, and obtaining one or more second test sequencing data sets determined by re-sequencing the one or more test nucleic acid molecules using non-terminating nucleotides provided in separate flow positions according to a second flow-cycle order different from the first, wherein the first flow-cycle order and second flow-cycle order are different.
Likelihood-based match scoring to candidate sequences across data sets
Determining, for each first sequencing data set and corresponding second sequencing data set, a respective match score to one or more candidate sequences, wherein the respective match score is indicative of a likelihood that the first test sequencing data set, the second test sequencing data set, or both, matches a candidate sequence from the one or more candidate sequences.
Calling variant presence or absence using determined match scores
Calling, using the determined match scores, the presence or absence of the short genetic variant in the test sample.
Sequencing by primer extension with labeled non-terminating nucleotides
Hybridizing the nucleic acid molecule to a primer to form a hybridized template, extending the primer using labeled, non-terminating nucleotides provided in separate flow positions according to a repeated flow-cycle order comprising five or more separate nucleotide flows.
Detecting signal or absence during extension
Detecting a signal from an incorporated labeled nucleotide or an absence of a signal as the primer is extended by the nucleotide flows.
Across the independent claims, the document claims disease-associated detection of a selected short genetic variant by computing likelihood-indicative match scores in flow-position space and calling presence/absence. Additional coverage includes obtaining two sets of test sequencing data using different flow-cycle orders and using corresponding match scores to candidate sequences for calling. A further independent aspect claims sequencing by primer extension with labeled non-terminating nucleotides delivered in a repeated flow-cycle order with five or more separate flows, together with signal/absence detection as the primer is extended.
Stated Advantages
Not explicitly described in patent.
Documented Applications
Not explicitly described in patent.
Interested in licensing this patent?