Secure communication of sensitive genomic information using probabilistic data structures
Inventors
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Assignees
NoblisNoblis is a nonprofit research and technical organization supporting federal missions in defense, health, environment, and security. Emphasizing applied sciences, engineering, digital transformation, artificial intelligence, cloud, and cybersecurity, Noblis provides objective solutions for government agencies confronting complex operational and scientific challenges.
Noblis is a nonprofit research and technical organization supporting federal missions in defense, health, environment, and security. Emphasizing applied sciences, engineering, digital transformation, artificial intelligence, cloud, and cybersecurity, Noblis provides objective solutions for government agencies confronting complex operational and scientific challenges.
Abstract
Techniques for securely encoding, communicating, and comparing genomic information using probabilistic data structures are provided. In some embodiments, genomic information in a secure computing environment may be encoded and/or anonymized by building a probabilistic data structure that represents sub-strings of the genomic information as members of a set; the probabilistic data structure may then be securely transmitted outside the secure computing environment. In some embodiments, a probabilistic data structure representing sub-strings of sensitive genomic information as members of a set may be received in an unsecure computing environment and may be queried to generate output data indicating whether reference sub-strings are probable members of the set. In some embodiments, querying the probabilistic data structure, and other techniques of analyzing the probabilistic data structure, may be used to determine whether the sensitive genomic information corresponds to an organism associated with the reference genomic information.
Core Innovation
The invention encodes nucleic acid sequence data by dividing received data representing a nucleic acid sequence into a plurality of portions, where each portion represents a sub-string of the nucleic acid sequence. For one or more of the portions, the system generates data including a probabilistic data structure that represents each portion as members of a set. The probabilistic data structure has a false-positive probability with no false negatives, so membership is not lost while false positives may occur.
The encoding selects the false-positive probability of the probabilistic data structure based at least in part on available storage resources on a non-transitory storage medium. A first false-positive probability is selected when a first amount of storage resources is available, and a second false-positive probability is selected when a second amount of storage resources is available, with the first false-positive probability lower than the second false-positive probability when the first amount of storage resources is higher than the second amount of storage resources. The encoded nucleic acid sequence including the probabilistic data structure is stored on the non-transitory storage medium.
The invention also supports probabilistic querying and comparison of encoded genomic information. In an unsecure computing environment, reference k-mers are used to query the probabilistic data structure to generate probabilistic membership and coverage metrics, including determining likely organism, species, and strain matches. The description further includes comparing two probabilistic data structures directly to compute a similarity index, including a Jaccard index, and applying thresholds to classify matches, with data flow between secure and unsecure computing environments using encoded genomic information database and metadata.
Claims Coverage
The independent claims, three total, cover encoding nucleic acid sequence data into a probabilistic data structure on a non-transitory storage medium, with selection of a false-positive probability based on available storage resources. The claim set further extends coverage via dependent claims to membership querying outcomes and to basis selection for the false-positive probability using other factors such as sensitivity, target file size, available processing resources, and accuracy requirements for comparisons.
Encoding by dividing nucleic acid sequence into substrings and representing portions as set members using a probabilistic data structure
Receive data representing a nucleic acid sequence, divide the data into a plurality of portions each representing a sub-string, and encode the data by generating data including a probabilistic data structure that represents each of the one or more portions as members of a set.
Selecting false-positive probability based on available storage resources with storage-dependent first and second probabilities
Generate the probabilistic data structure by selecting a false-positive probability based at least in part on available storage resources on a non-transitory storage medium, selecting a first false-positive probability if a first amount of storage resources is available and selecting a second false-positive probability if a second amount of storage resources is available, wherein the first false-positive probability is lower than the second false-positive probability and wherein the first amount of storage resources is higher than the second amount of storage resources.
Storing encoded nucleic acid sequence including the probabilistic data structure on non-transitory storage
Store the encoded nucleic acid sequence, including the probabilistic data structure, on the non-transitory storage medium.
Storing instructions to perform probabilistic encoding with storage-dependent false-positive probability
Store instructions on a non-transitory computer-readable storage medium that, when executed, cause a system to receive nucleic acid sequence data, divide it into portions as sub-strings, generate a probabilistic data structure representing portions as set members, select a false-positive probability based at least in part on available storage resources with a first and second probability tied to first and second amounts of storage resources, and store the encoded nucleic acid sequence including the probabilistic data structure on the non-transitory storage medium.
Across the independent claims, the inventive concept centers on encoding nucleic acid sequence data by representing sub-string portions as members of a set using a probabilistic data structure, while choosing the probabilistic false-positive probability based on available storage resources and storing the resulting encoded nucleic acid sequence including the probabilistic data structure.
Stated Advantages
Not explicitly described in patent.
Documented Applications
Not explicitly described in patent.
Interested in licensing this patent?
