Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Assignees
NoblisNoblis is a nonprofit research and technical organization supporting federal missions in defense, health, environment, and security. Emphasizing applied sciences, engineering, digital transformation, artificial intelligence, cloud, and cybersecurity, Noblis provides objective solutions for government agencies confronting complex operational and scientific challenges.
Noblis is a nonprofit research and technical organization supporting federal missions in defense, health, environment, and security. Emphasizing applied sciences, engineering, digital transformation, artificial intelligence, cloud, and cybersecurity, Noblis provides objective solutions for government agencies confronting complex operational and scientific challenges.
Abstract
Techniques for identifying regions in nucleic acid sequences for which to design highly discriminatory primers are provided. In some embodiments, a corpus of nucleic acid sequences may be divided into a first set and a second set, and a respective index may be built containing data structures representing a plurality of k-mers of each nucleic acid sequence. By comparing the data structures of the first index to one another, a system may iteratively determine whether each k-mer over a given region in one of the nucleic acid sequences in the first set are also found in every other sequence in the first set. By comparing against the data structures in the second index, a system may then iteratively determine whether all k-mers in the region can be found in the same order of in any of the nucleic acid sequences in the second set.
Core Innovation
The method receives genomic data representing a plurality of nucleic acid sequences and partitions the sequences into a first set and a second set. It creates and stores data in a first index representing the first set and data in a second index representing the second set, where each index comprises at least 4^12 elements and each element represents a respective permutation of nucleic acid sequences.
A target region is identified for which to design one or more primers that selects for one or more nucleic acid sequences in the first set and discriminates against one or more nucleic acid sequences in the second set. The identifying includes identifying, by the first index, the target region as a conserved region appearing in every nucleic acid sequence in the first set, and confirming, by the second index, that the conserved region appears in none of the nucleic acid sequences in the second set.
The indexing-based conserved-region determination is performed by extracting sub-strings (k-mers) and checking conservation across the first set and non-occurrence against the second set using position/correspondence and order/adjacency/position criteria. Conserved signature regions suitable for highly discriminatory primers are obtained by requiring the conserved region to be present in every sequence in the first set while being absent from the second set according to the second index comparisons, and the workflow generates and outputs data representing the identified target region.
Claims Coverage
The independent claims in this partial record cover a computational method, a system, and a non-transitory computer-readable storage medium that identify a target region for primer design using two separate indices over a first set and a second set, with conservation confirmed by one index and non-presence confirmed by the other. Across the independent claims, the core inventive features are the two-index permutation-based representation, the conserved-region identification/confirmation criteria, and output of data representing the identified target region.
Two indexed permutation-based sequence sets for primer target identification
creating and storing data in a first index representing a first set of the plurality of nucleic acid sequences, wherein the first index comprises at least 4^12 elements, wherein each of the 4^12 elements represents a respective permutation of nucleic acid sequences, and wherein the data created and stored in the first index comprises a first plurality of data structures each associated with a respective nucleic acid sequence of the first set; creating and storing data in a second index representing a second set of the plurality of nucleic acid sequences, wherein the second index comprises at least 4^12 elements, wherein each of the 4^12 elements represents a respective permutation of nucleic acid sequences, and wherein the data created and stored in the second index comprises a second plurality of data structures each associated with a respective nucleic acid sequence of the second set.
Conserved region present in every first-set sequence and absent from every second-set sequence
identifying a target region for which to design a primer that selects for one or more of the nucleic acid sequences in the first set and that discriminates against one or more of the nucleic acid sequences in the second set, wherein the identifying comprises: identifying, by the first index, the target region as a conserved region appearing in every nucleic acid sequence in the first set; and confirming, by the second index, that the conserved region appears in none of the nucleic acid sequences in the second set.
Output data representing the identified target region for primer design
generating and outputting data representing the identified target region.
The independent claims collectively require a computational framework that builds a first index and a second index over a first set and a second set of genomic sequences using permutation-based index elements, identifies a target region that is conserved across every sequence in the first set, confirms that the same conserved region appears in none of the second-set sequences, and outputs data representing the identified target region.
Stated Advantages
Documented Applications
No documented applications found
Interested in licensing this patent?
