Machine-learning based query construction and pattern identification
Inventors
Shukla, Oodaye • Finkbiner, Amy • Lauer, Robert • Garges, Cody • Izmailov, Rauf • Chadha, Ritu • Chiang, Cho-Yu Jason
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
A method, computer program product, and system identifying a probability of a medical condition in a patient. The method includes a processor obtaining data set(s) related to a patient population diagnosed with a medical condition and based on a frequency of features in the data set(s), identifying common features. The processor generates pattern(s) including a portion of the common features to generate a machine learning algorithm(s). The processor compiles a training set of data to use to tune the machine learning algorithm(s). The processor dynamically adjusts common features in the pattern(s) such that the machine learning algorithm(s) can distinguish patient data indicating the medical condition from patient data not indicating the medical condition. The processor applies the machine learning algorithm(s) to data related to the undiagnosed patient, to determine the probability.
Core Innovation
The invention is a computer-implemented method in a distributed computing environment for predicting a medical condition that is an orphan disease in an undiagnosed patient. It obtains one or more machine-readable data sets related to a patient population diagnosed with an orphan disease from one or more databases, and identifies common features in the data sets based on a frequency of features.
The common features are selected from diagnoses, medications, provider visits, treatment locations, medical procedures, ICD-9 codes, ICD-10 codes, diagnosis codes, procedure codes, adjudicated claims, care providers, disorders, diseases, provider types, care facility types, patient demographics, and co-morbidities. The invention generates one or more patterns comprising a portion of the identified common features by comparing members of a general population comprising the patient population diagnosed with the orphan disease and additional members not in the patient population diagnosed with the medical condition.
The invention further generates one or more machine learning algorithms based on the patterns to identify presence or absence of the given medical condition in an undiagnosed patient, and trains the algorithms by applying statistical sampling to compile a training set. The trained system determines a probability that indicates a percentage of commonality between the data related to the undiagnosed patient and the patterns, and that probability indicates a probability that the undiagnosed patient will be diagnosed with the medical condition in the future.
Claims Coverage
The provided excerpt includes three independent claim types: a computer-implemented method, a computer program product, and a system. Across these independent claims, the inventive features collectively cover 7 features.
Distributed obtaining of orphan disease patient datasets
Obtaining, by one or more processors in a distributed computing environment, one or more machine-readable data sets related to a patient population diagnosed with a medical condition from one or more databases, wherein the medical condition is an orphan disease.
Frequency-based selection of common features using identification methods
Based on a frequency of features in the one or more data sets, identifying common features in the one or more data sets utilizing an identification method-selected from weighting the common features based on frequency of occurrence in the one or more data sets, performing diffusion mapping, performing principal component analysis, performing recursive feature elimination, and utilizing a Random Forest to select the features; wherein if the selected identification method is weighting the common features, the common features comprise mutual information; wherein the identified common features are selected from diagnoses, medications, provider visits, treatment locations, medical procedures, ICD-9 codes, ICD-10 codes, diagnosis codes, procedure codes, adjudicated claims, care providers, disorders, diseases, provider types, care facility types, patient demographics, and co-morbidities.
Differentiating pattern generation by general population comparison
Generating one or more patterns comprising a portion of the common features identified in the one or more machine-readable data sets related to the patient population diagnosed with the medical condition from the one or more databases, wherein the portion comprises a smallest number of the common features that is a largest number of differentiating characteristics of the patient population diagnosed with the medical condition, wherein the portion of the common features is identified based on comparing members of a general population comprising the patient population diagnosed with the medical condition and additional members not in the patient population diagnosed with the medical condition, wherein a portion of the patient population diagnosed with the medical condition has the differentiating characteristics, and wherein the portion comprises fewer patients than the patient population diagnosed with the medical condition.
Machine learning generation for presence or absence in an undiagnosed patient
Generating one or more machine learning algorithms based on the one or more patterns, the one or more machine learning algorithms to identify presence or absence of the given medical condition in an undiagnosed patient, wherein the generating the one or more machine learning algorithms is based on absence or presence of features comprising the one or more patterns in data related to the undiagnosed patient.
Training set compilation via statistical sampling with additional non-condition population
Utilizing statistical sampling to compile a training set of data, wherein the training set comprises data from the one or more data sets and at least one additional data set comprising data related to a population without the medical condition.
Distributed query execution based on single-value vs parallel distribution
Utilizing the statistical sampling comprises formulating and obtaining queries based on the data set and processing and responding to the queries, the processing comprising, for each query: evaluating the query to determine if a prospective response to the query is a single value pulled from a single data set; based on determining that the prospective response to the query is the single value pulled from the single data set, assigning the query to a given computing resource in the distributed computing environment; and based on determining that the prospective response to the query is not the single value pulled from the single data set, distributing the query over a group of computing resources of the distributed computing environment to maximize efficiency, wherein the distributing comprises assigning each computing resource of the group of computing resources a portion of the query to execute in parallel with at least one other computing resource of the group of computing resources executing another portion of the query.
Future-diagnosis probability output as percentage of commonality
Determining, based on applying the one or more machine learning algorithms to data related to the undiagnosed patient, a probability, wherein the probability is a numerical value indicating a percentage of commonality between the data related to the undiagnosed patient and the one or more patterns, wherein the probability indicates a probability that the undiagnosed patient will be diagnosed with the medical condition in the future.
Across the independent claims, the core coverage is a distributed workflow that selects common features from orphan-disease patient datasets using an identification method, generates differentiating patterns by comparing with a general population, trains machine learning algorithms using statistical sampling that includes a non-condition population and distributed query execution rules, and outputs a probability indicating future diagnosis likelihood as a percentage of commonality.
Stated Advantages
Real-time or near-real-time analysis.
Reduced computational inefficiency via distributed query execution based on anticipated query complexity.
Documented Applications
Predicting presence or absence of an orphan disease in an undiagnosed patient and providing a probability that the undiagnosed patient will be diagnosed with the medical condition in the future.
Interested in licensing this patent?