Primer design using indexed genomic information

Inventors

IVANCICH, Mychal W.MONTOYA, Danielle E.

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Assignees

Noblis Inc

Member
Noblis
Noblis

Noblis is a nonprofit research and technical organization supporting federal missions in defense, health, environment, and security. Emphasizing applied sciences, engineering, digital transformation, artificial intelligence, cloud, and cybersecurity, Noblis provides objective solutions for government agencies confronting complex operational and scientific challenges.

Publication Number

US-11222712-B2

Patent

Publication Date

2022-01-11

Expiration Date


Abstract

Techniques for identifying regions in nucleic acid sequences for which to design highly discriminatory primers are provided. In some embodiments, a corpus of nucleic acid sequences may be divided into a first set and a second set, and a respective index may be built containing data structures representing a plurality of k-mers of each nucleic acid sequence. By comparing the data structures of the first index to one another, a system may iteratively determine whether each k-mer over a given region in one of the nucleic acid sequences in the first set are also found in every other sequence in the first set. By comparing against the data structures in the second index, a system may then iteratively determine whether all k-mers in the region can be found in the same order of in any of the nucleic acid sequences in the second set.

Core Innovation

The method receives genomic data representing a plurality of nucleic acid sequences and partitions the sequences into a first set and a second set. It creates and stores data in a first index representing the first set and data in a second index representing the second set, where each index comprises at least 4^12 elements and each element represents a respective permutation of nucleic acid sequences.

A target region is identified for which to design one or more primers that selects for one or more nucleic acid sequences in the first set and discriminates against one or more nucleic acid sequences in the second set. The identifying includes identifying, by the first index, the target region as a conserved region appearing in every nucleic acid sequence in the first set, and confirming, by the second index, that the conserved region appears in none of the nucleic acid sequences in the second set.

The indexing-based conserved-region determination is performed by extracting sub-strings (k-mers) and checking conservation across the first set and non-occurrence against the second set using position/correspondence and order/adjacency/position criteria. Conserved signature regions suitable for highly discriminatory primers are obtained by requiring the conserved region to be present in every sequence in the first set while being absent from the second set according to the second index comparisons, and the workflow generates and outputs data representing the identified target region.

Claims Coverage

The independent claims in this partial record cover a computational method, a system, and a non-transitory computer-readable storage medium that identify a target region for primer design using two separate indices over a first set and a second set, with conservation confirmed by one index and non-presence confirmed by the other. Across the independent claims, the core inventive features are the two-index permutation-based representation, the conserved-region identification/confirmation criteria, and output of data representing the identified target region.

Two indexed permutation-based sequence sets for primer target identification

creating and storing data in a first index representing a first set of the plurality of nucleic acid sequences, wherein the first index comprises at least 4^12 elements, wherein each of the 4^12 elements represents a respective permutation of nucleic acid sequences, and wherein the data created and stored in the first index comprises a first plurality of data structures each associated with a respective nucleic acid sequence of the first set; creating and storing data in a second index representing a second set of the plurality of nucleic acid sequences, wherein the second index comprises at least 4^12 elements, wherein each of the 4^12 elements represents a respective permutation of nucleic acid sequences, and wherein the data created and stored in the second index comprises a second plurality of data structures each associated with a respective nucleic acid sequence of the second set.

Conserved region present in every first-set sequence and absent from every second-set sequence

identifying a target region for which to design a primer that selects for one or more of the nucleic acid sequences in the first set and that discriminates against one or more of the nucleic acid sequences in the second set, wherein the identifying comprises: identifying, by the first index, the target region as a conserved region appearing in every nucleic acid sequence in the first set; and confirming, by the second index, that the conserved region appears in none of the nucleic acid sequences in the second set.

Output data representing the identified target region for primer design

generating and outputting data representing the identified target region.

The independent claims collectively require a computational framework that builds a first index and a second index over a first set and a second set of genomic sequences using permutation-based index elements, identifies a target region that is conserved across every sequence in the first set, confirms that the same conserved region appears in none of the second-set sequences, and outputs data representing the identified target region.

Stated Advantages

Documented Applications

No documented applications found

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.