Encoding text into nucleic acid sequences
Inventors
Hutchison, III, Clyde A. • Montague, Michael G. • Smith, Hamilton O.
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
Methods and apparatus are disclosed herein for encoding human readable text conveying a non-genetic message into nucleic acid sequences with a substantially reduced probability of biological impact and decoding such text from nucleic acid sequences. In one embodiment, each symbol of a symbol set of human readable symbols uniquely maps to a respective codon identifier. Mapping may ensure that each symbol will not map to a codon identifier that generates an amino acid residue which has a single-letter abbreviation that is the equivalent to the respective symbol. Synthetic nucleic acid sequences comprising such human readable text, and recombinant or synthetic cells comprising such sequences are provided, as well as methods of identifying cells, organisms, or samples containing such sequences.
Core Innovation
The invention provides a method of generating a sequence of codon identifiers corresponding to a sequence of human readable symbols, where the codon identifiers are assigned according to a coding scheme to convey a non-genetic message in a human reference language. The method receives the sequence of human readable symbols at a memory module, loads a human readable symbol map configured to map human readable symbols to specific codon identifiers, and uses a transcoder to map codon identifiers corresponding to each human readable symbol and output the sequence. The method then synthesizes a nucleic acid with the mapped sequence.
The human readable symbol map translates between codon identifiers and human reference language symbols, including specific mappings for characters and whitespace such as “space” and “new line,” and includes mapping definitions that do not use codon-to-amino-acid-letter equivalents. The design principles include mapping infrequent symbols to start codons and frequent symbols to stop codons, and optionally using all-6 reading-frame stop sequences flanking the message to support decoding across reading frames.
The invention also encompasses apparatus and computer-readable medium implementations, including a processor and storage module that store a mapping data structure to convert a sequence of codon identifiers into a human readable symbol sequence. It further describes synthetic nucleic acid sequences and recombinant or synthetic organisms, cells, or viruses containing non-genetic messages for authentication and detection using designed watermark sequences. The decoding concept includes analyzing all reading frames during decoding and comparing the decoded symbol sequence to a reference watermark.
Claims Coverage
The provided claim text centers on transcoders that map human readable symbols to codon identifiers using a stored symbol map and then synthesize nucleic acids carrying the mapped codon identifiers as a non-genetic message. Coverage further includes watermark-based authentication content, watermark component types, positional all-6 reading-frame stop-codon flanking sequences, and an apparatus implementation using a processor, storage module, and mapping data structure.
Generating codon-identifier sequence from human readable symbols via a symbol map and transcoder
receiving the sequence of human readable symbols at a memory module; loading a human readable symbol map configured to map human readable symbols to codon identifiers; using a transcoder to map a sequence of codon identifiers corresponding to each human readable symbol within the sequence according to the human readable symbol map and outputting the sequence; synthesizing a nucleic acid with the sequence
Using a human-readable watermark in the conveyed non-genetic message
the method wherein the human readable symbols include a watermark configured to authenticate and identify the recombinant or synthetic organism
Watermark content comprising copyright/trademark and other identifiers
the watermark comprises one or more forms of copyright notice, trademark, company identifier, name, phrase, sentence, quotation, genetic information, unique identifying information, data, or combinations thereof
Flanking nucleic acid with an all-6 reading frame stop-codon-containing sequence
including synthetic nucleic acid sequences that contain all-6 reading frame stop codon sequences located 5′ to a first codon identifier and/or 3′ to the last codon identifier in the sequence
Apparatus for converting codon identifiers to human readable symbol sequence
an apparatus comprising a processor and a storage module with a codon-identifier-to-human-readable-symbol mapping data structure to convert a sequence of codon identifiers into a human readable nucleic-acid-like symbol sequence based on that mapping
All-6 reading frame stop codon sequences at the 5′ and/or 3′ ends of codon-identifier sequence
the apparatus wherein a codon-identifier sequence contains stop-codon sequences for an all-6 reading frame both at the 5′ end relative to the first codon identifier and/or at the 3′ end relative to the last codon identifier
The claim coverage centers on mapping human readable symbols to codon identifiers with a stored symbol map and transcoder, synthesizing nucleic acids carrying a non-genetic message, and adding watermark and all-6 reading-frame stop-codon features for authentication and decoding. It also covers an apparatus with a processor, storage module, and mapping data structure for converting codon identifiers back to human readable symbol sequences.
Stated Advantages
Documented Applications
Authentication and detection using recombinant or synthetic organisms, cells, or viruses containing non-genetic messages with designed watermark sequences.
Example implementation using a synthetic Mycoplasma genome containing designed watermark sequences, with reported screening, restriction analysis, sequencing match, and successful recovery.
Interested in licensing this patent?