Systems and methods for cohort analysis using compressed data objects enabling fast memory lookups
Inventors
Shah, Nigam H. • Polony, Vladimir • Banda, Juan Manuel • Callahan, Alison Victoria
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
Systems and methods for structuring unstructured data according to a data object structure that enables fast query look-ups across a variety of space and time dimensions. Furthermore, many embodiments optimize the storage of the data objects using a set of compression techniques that configure the data types used for the data objects based on properties of the stored data. Furthermore, many embodiments provide are able to service query look-up requests without having to deserialize data within the byte stream format as stored in memory by encoding information that provide memory locations for requested data, thereby allowing for the immediate retrieval of the data as it is stored in the persistent memory.
Core Innovation
The invention provides a cohort analysis system that receives a search query to analyze medical data related to a patient, where the search query includes a plurality of parameters related to data components for a plurality of patient demographics. For the search query, the system determines a data object that corresponds to the patient and that stores a particular data component relevant to the search query. The data object comprises a plurality of data components, where each data component corresponds to a different type of data value related to the medical data.
The medical data for the patient is encoded within the plurality of data components in a serialized in-memory byte-stream format, and the medical data is received from a plurality of different medical information sources. The data object is stored, in its entirety, at a unique and continuous memory location. At least one header provides information regarding memory mappings of the plurality of data components within a body of the data object, including an offset for each data component and encoding information that identifies at least one data type used in storing the offset.
When responding to the search query, the system retrieves a particular data value directly from the particular data component without deserializing the data object, while the data object remains serialized. The retrieval uses the encoding information, the memory mappings, and the offset to identify a memory location of the particular data component, and the retrieved data value is obtained in the serialized in-memory byte-stream format. The system then generates an identification of a cohort of patients for the search query, where the identification includes the particular data value.
Claims Coverage
The independent claims are directed to a method, a system, and a non-transitory computer-readable medium for computer-implemented medical data analysis and cohort generation. Across the independent claims, the coverage centers on parameterized patient-demographics search queries, serialized in-memory byte-stream patient data objects stored at unique and continuous memory locations, and direct retrieval of data values from serialized components using header-provided memory mappings, offsets, and encoding information without deserializing entire objects, followed by cohort identification that includes the retrieved value.
Parameterized patient-demographics search query and cohort identification
Receiving a search query to analyze medical data related to a patient, wherein the search query comprises a plurality of parameters related to data components for a plurality of patient demographics; and generating an identification of a cohort of patients for the search query, wherein the identification includes the particular data value.
Serialized in-memory byte-stream patient data objects with unique continuous memory location
Determining a data object storing a particular data component relevant to the search query, wherein the data object corresponds to the patient and comprises a plurality of data components, including the particular data component; encoding medical data for the patient within the plurality of data components in a serialized in-memory byte-stream format; receiving the medical data from a plurality of different medical information sources; and storing the data object, in its entirety, at a unique and continuous memory location.
Header-provided memory mappings using offsets and encoding information
Providing at least one header that provides information regarding memory mappings of the plurality of data components within a body of the data object, the at least one header comprising an offset for each of the plurality of data components in the body of the data object and encoding information, wherein the encoding information identifies at least one data type used in storing the offset for each of the plurality of data components.
Direct retrieval of serialized data values without deserializing entire object
Retrieving a particular data value, responding to the search query, directly from the particular data component by using the encoding information, the memory mappings, and the offset to identify a memory location of the particular data component, wherein the particular data value is retrieved in the serialized in-memory byte-stream format while the data object remains serialized.
The inventive coverage requires a patient-demographics parameterized search query that leads to cohort identification including a retrieved value; the patient medical data is represented as serialized in-memory byte-stream patient data objects stored at unique continuous memory locations; memory access is driven by at least one header containing offsets and encoding information; and the particular data value is retrieved directly from the serialized component while keeping the data object serialized.
Stated Advantages
Not explicitly described in patent.
Documented Applications
Not explicitly described in patent.
Interested in licensing this patent?