Systems and methods for removing human genetic data from genetic sequences
Inventors
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
Systems and methods for removing data from strings are disclosed. A system can access a first hash table that stores representations of a first set of strings, where the representations have predetermined number of characters. The system can generate a second hash table that stores a second representations of a string of a second set of strings, where the second representations have the predetermined number of characters. Upon determining that the first hash table includes at least one of the plurality of second representations of the string included in the second hash table, the system can increment a counter associated with the string. The system can generate a third set of strings by removing the string from the second set of strings responsive to determining that the counter satisfies a threshold, and transmit the third set of strings to a computing system.
Core Innovation
The invention provides a system that removes sensitive human-genome-matching sequence data from genetic strings prior to transmission. The system accesses a hash table that stores a plurality of first k-mers or representations of a first set of strings associated with a human genome or genetic data from a human genome corpus, and generates k-mers or fixed-length string segments from reads of clusters or a second set of strings obtained from a sample.
The system evaluates correspondence against the stored human-genome k-mers or representations using hash table matching. After determining matches, it determines a second number that matches at least one of the first k-mers in the hash table, and/or increments a counter associated with the cluster or with the second set of strings, and generates a subset by removing the cluster or the string responsive to the second number or the counter satisfying a threshold.
Only the filtered subset is transmitted to a computing system, so that human-genome-matching data is excluded from what is communicated. The disclosed workflow supports chromosome-specific hash tables, non-overlapping genome portions, and additional hash tables for second representations corresponding to genetic information obtained from a sample.
Claims Coverage
The independent claims cover hash-table-based identification of human-genome-matching k-mers or representations, counter-based decisioning using a threshold, removal of matching clusters or strings, and transmission of only the remaining subset. Across the independent claims, the inventive features include k-mer or representation hash tables for human genome and sample-derived data, correspondence determination, counter incrementing, removal responsive to threshold, and transmission of the filtered subset.
Hash-table access to human-genome first k-mers
Access a hash table that stores a plurality of first k-mers of a human genome, each first k-mer corresponding to a first number of characters (k).
Generate second k-mers from clusters and match against first k-mers
Generate a plurality of second k-mers of a read of a cluster of a plurality of clusters; determine a second number of the plurality of second k-mers that match at least one of the plurality of first k-mers in the hash table.
Remove clusters based on thresholded match count and transmit subset
Generate a subset of the plurality of clusters by removing the cluster responsive to the second number satisfying a threshold; transmit the subset of the plurality of clusters to a computing system.
Generate first and second hash tables with counter-associated clusters
Access a first hash table that stores a plurality of first k-mers of a human genome; generate a second hash table that stores the plurality of second k-mers of a read of a cluster of plurality of clusters obtained from a sample of a pathogen; responsive to determining that the first hash table includes at least one of the plurality of second k-mers of the read included in the second hash table, increment a counter associated with the cluster.
Remove clusters based on counter threshold and transmit subset
Generate a subset of the plurality of clusters by removing the cluster responsive to determining that the counter satisfies a threshold; transmit the subset of the plurality of clusters to a computing system.
Hash-table representations for a first set of strings and segment-based correspondence
Generate a hash table that stores a plurality of representations of a first set of strings, each representation corresponding to a predetermined number of characters; for each string in a second set of strings, generate a string segment of the string having the predetermined number of characters; responsive to determining that the string segment corresponds to at least one of the plurality of representations of the first set of strings in the hash table, increment a counter associated with the second set of strings.
Remove strings based on counter threshold
Remove the string from a plurality of strings responsive to the counter satisfying a threshold.
Human-genome representations matched via first and second hash tables with counter-associated second sets
Access a first hash table that stores a plurality of representations of a first set of strings, each representation corresponding to a predetermined number of characters; generate a second hash table that stores a plurality of second representations of a string of a second set of strings, each second representation corresponding to the predetermined number of characters; responsive to determining, based on the first hash table and the second hash table, that the first hash table includes at least one of the plurality of second representations of the string included in the second hash table, increment a counter associated with the second set of strings.
Generate third set by thresholded removal and transmit third set
Generate a third set of strings by removing the string from the second set of strings responsive to determining that the counter satisfies a threshold; transmit the third set of strings to a computing system.
Taken together, the independent claims define a system that uses hash tables of human-genome k-mers or representations to determine matches to sample-derived k-mers, segments, or strings, increments counters or counts matches, removes clusters or strings when a threshold is satisfied, and transmits only the remaining non-removed subset to another computing system.
Stated Advantages
Improved efficiency for large corpora compared with brute-force matching.
Documented Applications
Removing sensitive human-genome-matching sequence data from genetic strings prior to transmission to a remote computing system.
Filtering pathogen sample cluster or read data by removing clusters whose matches to human-genome k-mers meet or exceed a threshold before transmitting the filtered subset.
Interested in licensing this patent?