Automated database updating and curation
Inventors
Fry, Stephen • Ellis, Jeremy • Shabilla, Matthew
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
Systems and methods for retrieval of information from read-only databases that hold taxonomic-related and sequence-related data. A method may include receiving organism names from a taxonomy database and detecting new organism names. The method may also include retrieving hierarchical data and assigning the new organism names to buckets based on the hierarchical data. The method may further include receiving sequence data elements from a nucleotide database, identifying particular buckets to correspond to a screener data set, querying organism names assigned to the particular buckets with names of reference sequences of the sequence data elements, generating a mapping between the sequence data elements and organism names returned as a result of the queries, and storing the mapping.
Core Innovation
A method and related system operate on a taxonomy database and a locally managed database in which organism-name data are organized using a set of buckets, where each bucket corresponds to an organism category. The approach receives, from the taxonomy database managed within a first domain, a plurality of organism names and detects one or more new organism names that are not represented in the locally managed database. For each detected new organism name, the approach receives metadata associated with the new organism name, including category data representing one or more categories of organism types that characterize the organism named by the new organism name.
Using the received metadata, the method identifies hierarchical data indicating two or more positions within a taxonomy corresponding to the categories of organism types, and determines an extent to which the two or more positions are represented on a same branch within the taxonomy. The method then identifies a degree of continuity across organism categories corresponding to the two or more positions, where the degree of continuity is higher when the positions are represented on the same branch than when represented on different branches. Based on the degree of continuity, the method identifies a bucket of the set of buckets that corresponds to a same organism category as indicated by the new organism name, and assigns the new organism name to the identified bucket, where the bucket does not include the new organism name.
The method then receives, at the computing system, a set of sequence data elements from a nucleotide database managed within a second domain, where each sequence data element includes a reference sequence and a name of the reference sequence. The method identifies one or more particular buckets corresponding to a screener data set and, for each sequence data element, queries organism names assigned to the particular buckets using the name of the reference sequence to generate a mapping between data corresponding to the sequence data element and an organism name returned as a result of the query, and stores the mapping. The process is extended, in some aspects, to align reads and identify a detected reference sequence and detected organism name using the stored mapping.
Claims Coverage
The independent claims in the provided set are clm-00001, clm-00009, and clm-00016, each covering two inventive features: assigning new organism names to bucketed organism categories based on taxonomy metadata and a computed degree of continuity, and using selected buckets to query a nucleotide database and store mappings. These independent claims share the same core inventive structure, implemented as a method, a system, and a computer-program product.
Assigning new organism names to bucketed organism categories using degree of continuity from taxonomy branch representation
Detecting one or more new organism names not represented in a locally managed database where organism-name data are organized using a set of buckets corresponding to organism categories; receiving taxonomy metadata for the new organism names; identifying hierarchical data indicating two or more positions within a taxonomy; determining an extent of representation on a same branch; identifying a degree of continuity across organism categories where continuity is higher for same-branch representation; identifying a bucket corresponding to a same organism category based on the degree of continuity, and assigning the new organism name to the identified bucket that does not include the new organism name.
Querying nucleotide reference-sequence names using selected buckets to generate and store mappings
Receiving a set of sequence data elements from a nucleotide database where each element includes a reference sequence and a name of the reference sequence; identifying one or more particular buckets corresponding to a screener data set; for each sequence data element, querying organism names assigned to the particular buckets with the name of the reference sequence to generate a mapping between data corresponding to the sequence data element and an organism name returned as a result of the query; and storing the mapping.
System implementation of taxonomy-driven bucket assignment and nucleotide-query mapping
A system with data processors and non-transitory computer-readable storage medium instructions to perform the actions of receiving organism names from a taxonomy database within a first domain, detecting new organism names not represented in a locally managed database with bucketed organism-name data, receiving metadata including category data, identifying hierarchical data positions, determining same-branch extent, determining degree of continuity, identifying and assigning the bucket for each new organism name, receiving nucleotide database sequence data elements within a second domain, identifying particular buckets for a screener data set, querying by reference-sequence name, generating a mapping, and storing the mapping.
Computer-program product configured to perform taxonomy-driven bucket assignment and nucleotide-query mapping
A computer-program product tangibly embodied in a non-transitory machine-readable storage medium with instructions configured to perform receiving organism names from a taxonomy database within a first domain, detecting new organism names not represented in a locally managed bucket-organized database, receiving metadata with category data, identifying hierarchical data positions, determining same-branch extent, identifying degree of continuity, identifying and assigning the bucket for each new organism name to a bucket that does not include the new organism name, receiving nucleotide database sequence data elements within a second domain, identifying one or more particular buckets corresponding to a screener data set, querying using the name of the reference sequence, generating a mapping, and storing the mapping.
Across clm-00001, clm-00009, and clm-00016, the key inventive elements are assigning newly detected organism names to buckets based on a computed degree of continuity derived from taxonomy branch representation of metadata-indicated positions, and using selected buckets to query a nucleotide database by reference-sequence name to generate and store mappings between sequence data elements and organism names.
Stated Advantages
Documented Applications
No documented applications found
Interested in licensing this patent?