Synthetic nucleic acid spike-ins
Inventors
Christians, Fred C. • Vilfan, Igor D. • Kertesz, Michael • Blauwkamp, Timothy A. • Venkatasubrahmanyam, Shivkumar • Rosen, Michael • Sit, Rene
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
This disclosure provides methods for determining relative abundance of one or more non-host species in a sample from a host. Also provided are methods involving addition of known concentrations of synthetic nucleic acids to a sample and performing sequencing assays to identify non-host species such as pathogens. Also provided are methods of tracking samples, tracking reagents, and tracking diversity loss in sequencing assays.
Core Innovation
Synthetic nucleic acid spike-ins are provided for NGS-based pathogen/non-host quantification and quality control. The spike-ins comprise nucleic acid sequences that are not naturally occurring, and at least 1,000 of the spike-in nucleic acids comprise one or more identifying tag regions positioned downstream or upstream of a unique variable region. The unique variable region comprises at least five degenerate bases, and the sequences of the spike-ins are unique to each other.
The spike-ins are used to quantify target abundance while addressing diversity loss during sample processing and to support identification of individual spike-ins. The design supports universal normalization based on length-based recovery profiles and GC-content matching, and spike-in panels spanning length, GC-content, and melting temperature are used to infer denaturation/recovery behavior. Tracer/sample-identifier spike-ins are used for sample tracking and cross-contamination detection.
The document also describes designed spike-in chemistries to enable carrier effects while evading sequencing/library steps, including ligation resistance, end-repair resistance, and blocking groups or depletion/removal strategies. Molecular LIMS, sequencing-detectable spike-ins to barcode reagent lots and containers, and de-duplication based on identifying tag regions are also described.
Claims Coverage
The provided independent claims cover at least 1,000 synthetic spike-in or synthetic nucleic acid sequences that are not naturally occurring and are uniquely identifiable using identifying tag regions around a unique variable region with at least five degenerate bases. Additional independent claim coverage includes compositions defined by SEQ ID NO: 119 or SEQ ID NO: 120, compositions with a biological sample or genomic nucleic acids not attached to an adapter, and one claim with a specific grouping arrangement of the degenerate bases.
Non-naturally occurring synthetic spike-ins with unique identifying tag regions
A composition comprising at least 1,000 synthetic spike-in nucleic acids having nucleic acid sequences that are not naturally occurring, where at least 1,000 of the synthetic spike-in nucleic acids comprise one or more identifying tag regions positioned downstream or upstream of a unique variable region, the unique variable region comprises at least five degenerate bases, and sequences of the at least 1,000 synthetic spike-in nucleic acids are unique to each other.
Fixed-sequence synthetic nucleic acid identifiers
A composition comprising at least 1,000 unique synthetic nucleic acid sequences wherein each of the at least 1,000 unique synthetic nucleic acids comprise sequences identified in SEQ ID NO: 119 or SEQ ID NO: 120.
Biological sample composition with unique tagged degenerate-variable spike-ins
A composition comprising a biological sample from a tissue or body fluid and at least 1,000 synthetic nucleic acids comprising nucleic acid sequences that are not naturally occurring, wherein at least 1,000 of the synthetic nucleic acids comprise one or more identifying tag regions positioned downstream or upstream of a unique variable region, the unique variable region comprises at least five degenerate bases, and sequences of the at least 1,000 synthetic nucleic acids are unique to each other.
Genomic nucleic acids without adapter plus unique tagged degenerate-variable spike-ins
A composition comprising genomic nucleic acids that are not attached to an adapter and at least 1,000 synthetic nucleic acids comprising nucleic acid sequences that are not naturally occurring, wherein at least 1,000 of the synthetic nucleic acids comprise one or more identifying tag regions positioned downstream or upstream of a unique variable region, the unique variable region comprises at least five degenerate bases, and sequences of the at least 1,000 synthetic nucleic acid sequences are unique to each other.
Degenerate-variable region grouped into separated degenerate-base groups
A composition comprising at least 1,000 synthetic spike-in nucleic acids comprising non-natural nucleic acid sequences that are not naturally occurring, wherein at least 1,000 unique synthetic nucleic acids have one or more identifying tag regions positioned downstream or upstream of a unique variable region, the unique variable region comprises at least five degenerate bases, the degenerate bases are split into at least two groups separated by at least four contiguous nucleotides, and each non-natural nucleic acid sequence that is not naturally occurring is unique to each other.
Across the independent claims, the compositions are defined around at least 1,000 non-naturally occurring synthetic nucleic acids or spike-ins with unique identities. Uniqueness is supported by identifying tag regions flanking a unique variable region with at least five degenerate bases, with an additional structural arrangement requirement in the degenerate-base grouping claim and a fixed SEQ ID limitation in the SEQ ID composition claim.
Stated Advantages
Provides NGS-based pathogen/non-host quantification and quality control using synthetic spike-ins.
Measures diversity loss during sample processing to support computation of target abundance.
Supports universal normalization including length-based recovery profiles and GC-content matching.
Enables sample tracking and cross-contamination detection using tracer/sample-identifier spike-ins.
Infers denaturation/recovery behavior using spike-in panels spanning length, GC-content, and melting temperature.
Enables carrier effects while evading sequencing/library steps, including ligation resistance and end-repair resistance.
Enables quantification of recovery and denaturation effects.
Enables quantification via diversity loss and recovery.
Detects cross-contamination using tracer sequences in positive controls.
Supports barcoding of reagent lots/containers using sequencing-detectable spike-ins.
Supports identification of tag regions for sequencing analysis and de-duplication/deduping.
Documented Applications
NGS-based pathogen/non-host quantification and quality control using synthetic nucleic acid spike-ins.
Sample tracking and cross-contamination detection using tracer/sample-identifier spike-ins.
Assay context including use in sequencing/library preparation with sample types such as plasma, serum, CSF, urine, and other body fluids.
Infection detection/monitoring using the described spike-ins and sequencing-based quantification.
Cancer/disease diagnostics using the described spike-in and sequencing approaches.
Interested in licensing this patent?