Homopolymer encoded nucleic acid memory

Inventors

Efcavitch, J. WilliamHolden, Matthew T.

Assignees

Molecular Assemblies Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-11174512-B2

Patent

Publication Date

2021-11-16

Expiration Date


Abstract

Nucleic acid memory strands encoding digital data using a sequence of homopolymer tracts of repeated nucleotides provides a cheaper and faster alternative to conventional digital DNA storage techniques. The use of homopolymer tracts allows for lower fidelity, high throughput sequencing techniques such as nanopore sequencing to read data encoded in the memory strands. Specialized synthesis techniques allow for synthesis of long memory strands capable of encoding large volumes of data despite the reduced data density afforded by homopolymer tracts as compared to conventional single nucleotide sequences.

Core Innovation

The invention relates to recording data using a nucleic acid memory strand composed of homopolymer tracts or heteropolymer tracts. A dataset is represented as an in silico sequence of bits, and the nucleic acid memory strand is synthesized so that each bit is encoded in a homopolymer or heteropolymer tract.

For homopolymer tracts, bits are converted by identifying transitions between homopolymer tracts. For heteropolymer tracts, bits are encoded based on a ratio of two or more different nucleotides or nucleotide analogs within each heteropolymer tract.

The disclosed approach supports base 2, base 3, and base 4, and extension to higher bases using modified nucleotide or non-nucleotide analogs for multi-bit capacity. Template-independent synthesis is used to form tract length-defined memory strands, and nanopore readout enhancements include nanopore stoppers, hairpins, and trapped or circularized nanopore constructs for repeated reads.

Data protection is addressed using write-once read-many and encryption or access conditions using cleavable linkers and modifications.

Claims Coverage

The independent claim covers a nucleic acid data recording method with two main encoding mechanisms: homopolymer tract transitions and heteropolymer tract nucleotide or analog ratios.

In silico bit sequence representing a dataset

Creating an in silico sequence of bits that represents a dataset.

Synthesizing homopolymer or heteropolymer memory strands for bit encoding

Synthesizing a nucleic acid memory strand comprising a plurality of homopolymer tracts or heteropolymer tracts, wherein each bit of the sequence of bits representing the dataset is encoded in a homopolymer or heteropolymer tract; the heteropolymer tracts comprise two or more different nucleotides or nucleotide analogs.

Converting homopolymer tracts to bits using transitions

Converting homopolymer tracts to bits by identifying transitions between homopolymer tracts.

Encoding heteropolymer tract bits using nucleotide or analog ratios

Encoding bits for heteropolymer tracts based on a ratio of the two or more different nucleotides or nucleotide analogs to each other within each heteropolymer tract.

Claim coverage centers on encoding each dataset bit in a homopolymer tract via transitions or in a heteropolymer tract via nucleotide or analog ratios, using an in silico bit sequence to define the dataset representation.

Stated Advantages

Enables data storage and readout that relies on tract transitions and ratios rather than nucleotide-by-nucleotide fidelity, supporting higher-throughput or lower-fidelity readout.

Supports multi-bit capacity by using multiple bases and higher-base extensions via modified nucleotide or non-nucleotide analogs.

Supports longer memory strands by using template-independent synthesis and controlling tract length.

Supports repeated reads and improved nanopore readout via nanopore stoppers, hairpins, and trapped or circularized nanopore constructs.

Provides write-once read-many and encryption or access conditions using cleavable linkers and modifications.

Documented Applications

Nanopore readout of nucleic acid digital data storage based on homopolymer tract transitions and heteropolymer tract ratios.

Long-read sequencing contexts using long DNA strands for memory storage and readout.

WORM and encryption or access-condition data protection for the nucleic acid memory strands.

Encoding a text message into a base-two nucleotide code for homopolymer strand storage and subsequent readout.

Example workflow including encoding and readout using NGS library preparation and analysis.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.