Methods for detecting nucleic acid variants
Inventors
Almogy, Gilad • Pratt, Mark • BARAD, Omer • Faigler, Simchon • Oberstrass, Florian
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
Methods for detecting a short genetic variant in a test sample are described herein. In some exemplary methods, the short genetic variant is called using one or match scores, which are determined using one or more sequencing data sets obtained from a test nucleic acid molecule, wherein the test sequencing data sets are determined by sequencing the test nucleic acid molecule using non-terminating nucleotides provided in separate nucleotide flows according to a flow-cycle order. Also described herein are methods of sequencing a test nucleic acid molecule using two or more different flow-cycle orders and/or extended flow cycle orders having five or more nucleotide flows per flow cycle.
Core Innovation
The disclosure describes detecting a disease based on single nucleotide variants identified from sequencing by selecting a set of SNV loci from a disease-associated SNV locus panel generated by sequencing a nucleic acid sample derived from diseased tissue from a subject. The selected SNV loci are associated with a diseased sequencing data set that differs from a reference sequencing data set across at least one flow cycle when the diseased sequencing data set and the reference sequencing data set are obtained by sequencing using non-terminating nucleotides provided in separate flow positions according to a flow-cycle order.
The approach obtains a cell-free nucleic acid data set by sequencing a cell-free nucleic acid sample from the subject using non-terminating nucleotides provided in separate flow positions according to a flow-cycle order, where the mean sequencing depth of the cell-free nucleic acid sequencing data set is less than 10. A fraction value is then determined by processing a total number of SNV reads detected at the set of SNV loci, a number of loci selected in the set of SNV loci, the mean sequencing depth, and a sequencing false positive error rate, and the fraction value is used for calling disease presence, absence, progression, or regression.
The disclosure grounds the variant and disease signal in how sequencing data differ across flow cycles, including sequencing using extended flow cycle orders with at least five flows per cycle and determining target variants or candidates based on how many flow positions differ between target and reference datasets. The match scores are computed for candidate sequences using flow-space information derived from separate nucleotide flows, enabling presence/absence calling for short genetic variants such as SNPs or indels without computationally expensive alignment.
Claims Coverage
The partial content provides one independent claim directed to detecting a disease based on SNVs from cell-free nucleic acid sequencing using flow-cycle-based SNV locus selection and a fraction value that accounts for sequencing depth and a sequencing false positive error rate. Dependent claim text is not provided in the partial content, so only the independent claim can be extracted explicitly.
Flow-cycle-based selection of disease-associated SNV loci
Selecting a set of SNV loci from a disease-associated SNV locus panel, wherein the SNV loci in the selected set are associated with a diseased sequencing data set that differs from a reference sequencing data set across at least one flow cycle when both are obtained by sequencing using non-terminating nucleotides provided in separate flow positions according to a flow-cycle order.
Cell-free nucleic acid sequencing using non-terminating nucleotides in separate flow positions
Sequencing a cell-free nucleic acid sample from the subject using non-terminating nucleotides provided in separate flow positions according to a flow-cycle order to obtain a cell-free nucleic acid data set, wherein the mean sequencing depth of the cell-free nucleic acid sequencing data set is less than 10.
Fraction value using SNV read counts, loci number, mean depth, and sequencing false positive error rate
Determining a fraction value by processing a total number of SNV reads detected at the set of SNV loci in the cell-free nucleic acid data set, a number of loci selected in the set of SNV loci, a mean sequencing depth of the cell-free nucleic acid data set, and a sequencing false positive error rate.
Calling disease presence, absence, progression, or regression based on fraction value or its change
Calling a presence, absence, progression, or regression of the disease in the subject based on the fraction value or a degree of change in the fraction value from a prior fraction value determined for the subject.
The inventive coverage combines flow-cycle-based SNV locus selection from a disease-associated panel, low-depth cell-free nucleic acid sequencing using non-terminating nucleotides in separate flow positions, computation of a fraction value that incorporates sequencing false positive error rate, and disease calling from the fraction value or its change.
Stated Advantages
Supports residual disease detection and monitoring by calling presence, absence, progression, or regression based on a computed fraction value and/or its change.
Uses sequencing false positive error rate and mean sequencing depth within the fraction value determination to model noise and error in calling disease.
Enables presence/absence calling for short genetic variants such as SNPs or indels without computationally expensive alignment.
Documented Applications
Residual disease detection and monitoring using personalized SNV locus panels, including monitoring recurrence, progression, or regression based on changes in the fraction value.
Detecting a disease in a subject using cell-free nucleic acid sequencing based on selected SNV loci associated with diseased and reference sequencing data sets across flow cycles.
Presence/absence calling for short genetic variants such as SNPs or indels.
Interested in licensing this patent?