Systems and methods for sequencing error correction via double strand preservation

Inventors

Mazur, Daniel

Assignees

Ultima Genomics Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-12195797-B2

Patent

Publication Date

2025-01-14

Expiration Date


Abstract

The present disclosure provides compositions, methods, and systems for preparing nucleic acid samples for sequencing, including attaching adapters to a template nucleic acid molecule having a first and second strands; attaching the template nucleic acid molecule to a support; detaching the first and second strands from each other; generating a reverse complement copy of the second strand using a primer coupled to the support; generating amplified copies of the first strand and the reverse complement copy of the second strand using primers coupled to the support; and identifying the sequences of the first strand and the reverse complement copy of the second strand. Further provided herein are methods of error correction by preserving both strands of a template nucleic acid molecule during amplification.

Core Innovation

The invention relates to sequencing error correction by preserving sequences of both strands of a nucleic acid molecule during amplification. A double-stranded nucleic acid molecule is attached to a support that comprises a plurality of surface primers, with a first strand attached to a first surface primer and a second strand at least partially hybridized to the first strand. The method detaches the second strand, or a derivative thereof, from the first strand and anneals the second strand, or the derivative thereof, to a second surface primer.

The second surface primer is extended using the second strand, or the derivative thereof, as a template to generate an extended second surface primer comprising a reverse complement copy of the second strand. This generates an amplified support attached to both the first strand and the extended second surface primer comprising the reverse complement copy. In additional aspects, the second strand, or a derivative thereof, is generated by cleaving at a cleavage site adjacent to a 3′ end to produce a 3′-hydroxyl terminus, followed by extending using the first strand as a template.

Specific cleavage options include cleaving based on a uracil residue using Uracil-DNA glycosylase and Apurinic/apyrimidinic Endonuclease 1, or cleaving based on a ribonucleotide using RNase HII and Apurinic/apyrimidinic Endonuclease 1 enzymes. The resulting amplified support preserves sequence information from both strands while enabling an error-correction workflow tied to sequencing outputs.

Claims Coverage

The claims center on support-based amplification that preserves both strand sequences by detaching and re-annealing the second strand to a different surface primer and extending to produce a reverse-complement copy on the amplified support. Additional described refinements include 3′-adjacent cleavage to generate a second-strand derivative, use of an artificial base mismatch error for error identification in sequencing data, locus-level labeling of error SNPs and filtering from true SNPs, and a quantitative constraint on the first-strand copy ratio. There are eight inventive features.

Preserving both strands on a support during amplification

Preserving sequences of both strands of a nucleic acid molecule during amplification by attaching a double-stranded nucleic acid molecule to a support comprising a plurality of surface primers, where the first strand is attached to a first surface primer and the second strand is at least partially hybridized to the first strand.

Detaching and re-annealing the second strand to a different surface primer

Detaching the second strand, or derivative thereof, from the first strand and annealing the second strand, or the derivative thereof, to a second surface primer of the plurality of surface primers.

Extending to generate a reverse-complement copy on the amplified support

Extending the second surface primer using the second strand, or the derivative thereof, as a template to generate an extended second surface primer comprising a reverse complement copy of the second strand, thereby generating an amplified support attached to both the first strand and the extended second surface primer comprising the reverse complement copy.

Deriving the second-strand derivative via 3′-adjacent cleavage

Generating the second strand, or derivative thereof, by cleaving at a cleavage site adjacent to a 3′ end to produce a 3′-hydroxyl terminus and extending the second strand from the 3′-hydroxyl terminus using the first strand as a template to generate an extended second strand.

Using UDG and APE1 or RNase HII and APE1 for cleavage

Configuring the cleavage site to comprise either a uracil residue cleaved using Uracil-DNA glycosylase and Apurinic/apyrimidinic Endonuclease 1, or a ribonucleotide cleaved using RNase HII and APE1.

Introducing and identifying an artificial base mismatch error

Including an artificial base mismatch error that is not present in a native template nucleic acid molecule derived from the double-stranded nucleic acid molecule, and identifying the artificial base mismatch error from a sequencing dataset.

Classifying the mismatch locus as an error SNP and filtering

Identifying a locus of the artificial base mismatch error as an error single nucleotide polymorphism (SNP), and filtering error SNPs from true SNPs.

Restraining strand copy ratio on the amplified support

Setting a ratio of the plurality of copies of the first strand in the plurality of nucleic acid molecules on the amplified support to be between about 0.4 and 0.6.

The claims emphasize support-based amplification that preserves both strands by detaching and re-annealing the second strand to a different surface primer and extending to produce a reverse-complement copy on the amplified support. Additional described refinements include 3′-adjacent cleavage to generate a second-strand derivative, use of an artificial base mismatch error for error identification in sequencing data, locus-level labeling of error SNPs and filtering from true SNPs, and a quantitative constraint on the first-strand copy ratio.

Stated Advantages

Enables sequencing error correction by preserving sequences of both strands during amplification.

Improves error-locus handling in sequencing data analysis by identifying artificial base mismatch error loci that create phasing events leading to sequencing quality degradation and downstream trimming that removes error-locus signals.

Uses a computed read quality metric and trimming thresholds, including optional moving averages, to trim read termini based on error-related phasing and signals.

The resulting amplified support preserves sequence information from both strands while enabling an error-correction workflow tied to sequencing outputs.

Documented Applications

Flow-based sequencing signal interpretation using per-flow homopolymer length likelihoods (flowgram matrix), including phasing events and trimming based on read quality metric and thresholds.

Sequencing error correction.

Error-correction workflow tied to sequencing outputs.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.