Systems and methods for sequencing error correction via double strand preservation
Inventors
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
The present disclosure provides compositions, methods, and systems for preparing nucleic acid samples for sequencing, including attaching adapters to a template nucleic acid molecule having a first and second strands; attaching the template nucleic acid molecule to a support; detaching the first and second strands from each other; generating a reverse complement copy of the second strand using a primer coupled to the support; generating amplified copies of the first strand and the reverse complement copy of the second strand using primers coupled to the support; and identifying the sequences of the first strand and the reverse complement copy of the second strand. Further provided herein are methods of error correction by preserving both strands of a template nucleic acid molecule during amplification.
Core Innovation
The invention relates to sequencing error correction by preserving sequences of both strands of a nucleic acid molecule during amplification. A double-stranded nucleic acid molecule is attached to a support that comprises a plurality of surface primers, with a first strand attached to a first surface primer and a second strand at least partially hybridized to the first strand. The method detaches the second strand, or a derivative thereof, from the first strand and anneals the second strand, or the derivative thereof, to a second surface primer.
The second surface primer is extended using the second strand, or the derivative thereof, as a template to generate an extended second surface primer comprising a reverse complement copy of the second strand. This generates an amplified support attached to both the first strand and the extended second surface primer comprising the reverse complement copy. In additional aspects, the second strand, or a derivative thereof, is generated by cleaving at a cleavage site adjacent to a 3′ end to produce a 3′-hydroxyl terminus, followed by extending using the first strand as a template.
Specific cleavage options include cleaving based on a uracil residue using Uracil-DNA glycosylase and Apurinic/apyrimidinic Endonuclease 1, or cleaving based on a ribonucleotide using RNase HII and Apurinic/apyrimidinic Endonuclease 1 enzymes. The resulting amplified support preserves sequence information from both strands while enabling an error-correction workflow tied to sequencing outputs.
Claims Coverage
The claims center on support-based amplification that preserves both strand sequences by detaching and re-annealing the second strand to a different surface primer and extending to produce a reverse-complement copy on the amplified support. Additional described refinements include 3′-adjacent cleavage to generate a second-strand derivative, use of an artificial base mismatch error for error identification in sequencing data, locus-level labeling of error SNPs and filtering from true SNPs, and a quantitative constraint on the first-strand copy ratio. There are eight inventive features.
Preserving both strands on a support during amplification
Preserving sequences of both strands of a nucleic acid molecule during amplification by attaching a double-stranded nucleic acid molecule to a support comprising a plurality of surface primers, where the first strand is attached to a first surface primer and the second strand is at least partially hybridized to the first strand.
Detaching and re-annealing the second strand to a different surface primer
Detaching the second strand, or derivative thereof, from the first strand and annealing the second strand, or the derivative thereof, to a second surface primer of the plurality of surface primers.
Extending to generate a reverse-complement copy on the amplified support
Extending the second surface primer using the second strand, or the derivative thereof, as a template to generate an extended second surface primer comprising a reverse complement copy of the second strand, thereby generating an amplified support attached to both the first strand and the extended second surface primer comprising the reverse complement copy.
Deriving the second-strand derivative via 3′-adjacent cleavage
Generating the second strand, or derivative thereof, by cleaving at a cleavage site adjacent to a 3′ end to produce a 3′-hydroxyl terminus and extending the second strand from the 3′-hydroxyl terminus using the first strand as a template to generate an extended second strand.
Using UDG and APE1 or RNase HII and APE1 for cleavage
Configuring the cleavage site to comprise either a uracil residue cleaved using Uracil-DNA glycosylase and Apurinic/apyrimidinic Endonuclease 1, or a ribonucleotide cleaved using RNase HII and APE1.
Introducing and identifying an artificial base mismatch error
Including an artificial base mismatch error that is not present in a native template nucleic acid molecule derived from the double-stranded nucleic acid molecule, and identifying the artificial base mismatch error from a sequencing dataset.
Classifying the mismatch locus as an error SNP and filtering
Identifying a locus of the artificial base mismatch error as an error single nucleotide polymorphism (SNP), and filtering error SNPs from true SNPs.
Restraining strand copy ratio on the amplified support
Setting a ratio of the plurality of copies of the first strand in the plurality of nucleic acid molecules on the amplified support to be between about 0.4 and 0.6.
The claims emphasize support-based amplification that preserves both strands by detaching and re-annealing the second strand to a different surface primer and extending to produce a reverse-complement copy on the amplified support. Additional described refinements include 3′-adjacent cleavage to generate a second-strand derivative, use of an artificial base mismatch error for error identification in sequencing data, locus-level labeling of error SNPs and filtering from true SNPs, and a quantitative constraint on the first-strand copy ratio.
Stated Advantages
Enables sequencing error correction by preserving sequences of both strands during amplification.
Improves error-locus handling in sequencing data analysis by identifying artificial base mismatch error loci that create phasing events leading to sequencing quality degradation and downstream trimming that removes error-locus signals.
Uses a computed read quality metric and trimming thresholds, including optional moving averages, to trim read termini based on error-related phasing and signals.
The resulting amplified support preserves sequence information from both strands while enabling an error-correction workflow tied to sequencing outputs.
Documented Applications
Flow-based sequencing signal interpretation using per-flow homopolymer length likelihoods (flowgram matrix), including phasing events and trimming based on read quality metric and thresholds.
Sequencing error correction.
Error-correction workflow tied to sequencing outputs.
Interested in licensing this patent?