Materials and methods for the synthesis of error-minimized nucleic acid molecules

Inventors

Gibson, Daniel G. • Caiazza, Nicky • Richardson, Toby H.

Assignees

Telesis Bio Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-10704041-B2

Patent

Publication Date

2020-07-07

Expiration Date


Abstract

The present invention provides materials and methods useful for error correction of nucleic acid molecules. In one embodiment of the invention, a first plurality of double-stranded nucleic acid molecules having a nucleotide mismatch are fragmented by exposure to a molecule having unidirectional mismatch endonuclease activity. The nucleic acid molecules are cut at the mismatch site or near the mismatch site, leaving a double-stranded nucleic acid molecule having a mismatch at the end or near end of the molecule. The nucleic acid molecule is then exposed to a molecule having unidirectional exonuclease activity to remove the mismatched nucleotide. The missing nucleotides can then be filled in by the action of, e.g., a molecule having DNA polymerase activity. The result is double-stranded nucleic acid molecules with a decreased frequency of nucleotide mismatches. Also provided are novel nucleic acid sequences encoding mismatch endonucleases, polypeptides encoded thereby, as well as nucleic acid constructs, transgenic cells, and various compositions thereof.

Core Innovation

The disclosed invention relates to an error-correction strategy for synthetic/double-stranded nucleic acid molecules that reduces mismatches. The approach generates mismatches in double-stranded nucleic acid material, fragments the resulting double-stranded nucleic acids at mismatch sites using a unidirectional mismatch endonuclease activity, and removes the mismatches using a same-direction unidirectional exonuclease activity. The resulting processed nucleic acid material is filled/assembled to produce double-stranded DNA with a decreased mismatch frequency.

The invention further provides isolated mismatch-endonuclease genes and polypeptides, including mismatch endonuclease activities based on CEL I/CEL II/RES I/SURVEYOR-based nuclease families. Chimeric endonucleases are disclosed, including CEL I and CEL II mature cores, as well as core enzymes expressed in recombinant host cells. The document also discloses recombinant nucleic acid constructs that include heterologous nucleic acid linkage, including heterologous transcription control elements and heterologous linkage sequences.

Recombinant host cells and compositions/kits are disclosed for producing the mismatch endonuclease enzymes and for implementing the mismatch removal workflow concept. Experimental examples describe synthetic HA/NA gene error correction using endonuclease plus exonuclease workflows, including two-step and one-step endonuclease+exonuclease workflows. The examples report improved error rates and higher fractions of correct gene clones in assembled synthetic HA/NA genes.

Claims Coverage

The independent claim covers a recombinant nucleic acid molecule defined by sequence-identity thresholds to specified nucleic acid sequences or encoded polypeptides, together with an operable linkage to a heterologous nucleic acid. The claim family includes dependent claims that refine the sequence-identity thresholds, restrict the SEQ ID targets, and narrow to particular constructs and host systems.

Recombinant nucleic acid with SEQ ID identity thresholds

A recombinant nucleic acid molecule comprising a nucleic acid sequence exhibiting 70% or greater sequence identity to a nucleic acid sequence selected from the group consisting of SEQ ID NO: 09, SEQ ID NO: 12, SEQ ID NO: 15, SEQ ID NO: 18, SEQ ID NO: 20, SEQ ID NO: 22, SEQ ID NO: 24, SEQ ID NO: 26, SEQ ID NO: 30, SEQ ID NO: 32, a complement thereof or a fragment of either; or a nucleic acid sequence encoding a polypeptide exhibiting 80% or greater sequence identity to an amino acid sequence selected from the group consisting of SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 19, SEQ ID NO: 21, SEQ ID NO: 23, SEQ ID NO: 25, SEQ ID NO: 27, SEQ ID NO: 28, and SEQ ID NO: 29.

Operably linked to a heterologous nucleic acid

The nucleic acid sequence is operably linked to a heterologous nucleic acid.

Higher sequence-identity recombinant nucleic acid

A recombinant nucleic acid molecule comprising a nucleic acid sequence exhibiting 90% or greater sequence identity to a nucleic acid sequence selected from the group consisting of SEQ ID NO: 09, SEQ ID NO: 12, SEQ ID NO: 15, SEQ ID NO: 18, SEQ ID NO: 20, SEQ ID NO: 22, SEQ ID NO: 24, SEQ ID NO: 26, SEQ ID NO: 30, and SEQ ID NO: 32, a complement thereof or a fragment of either; or a nucleic acid sequence encoding a polypeptide exhibiting 90% or greater sequence identity to an amino acid sequence selected from the group consisting of SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 19, SEQ ID NO: 21, SEQ ID NO: 23, SEQ ID NO: 25, SEQ ID NO: 27, SEQ ID NO: 28, and SEQ ID NO: 29.

Further tightened polypeptide identity to specified SEQ IDs

A recombinant nucleic acid molecule comprises a nucleic acid sequence encoding a polypeptide that has at least 95% identity to an amino acid sequence selected from SEQ ID NO: 10 or SEQ ID NO: 16.

Encoded mismatch endonuclease activity

A recombinant nucleic acid molecule including a nucleic acid sequence encoding a molecule that has mismatch endonuclease activity.

Recombinant host cell containing the nucleic acid construct

A recombinant host cell includes a nucleic acid construct comprising a nucleic acid molecule as defined in claim 2.

Isolated mismatch-endonuclease polypeptide with defined residue ranges

The isolated polypeptide comprises an amino-acid sequence chosen from specified SEQ ID numbers and defined residue ranges from those SEQ ID numbers.

Across the independent claim and its dependent refinements, the claimed subject matter is centered on a recombinant nucleic acid molecule defined by specified SEQ ID-based nucleic-acid and/or polypeptide identity thresholds and the requirement that the nucleic acid sequence is operably linked to a heterologous nucleic acid, with further dependent limitations directed to mismatch endonuclease activity, recombinant host cell embodiments, and isolated polypeptide residue-range variants.

Stated Advantages

Improved error rates.

Higher fractions of correct gene clones in assembled synthetic HA/NA genes.

Decreased mismatch frequency.

Documented Applications

Synthetic HA/NA gene error correction.

Producing mismatch endonuclease enzymes.

Implementing the mismatch removal workflow concept.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.