Machine-learning based query construction and pattern identification for amyotrophic lateral sclerosis
Inventors
Miller, Chris • Kasoji, Manjula • Shukla, Oodaye • Garges, Cody • Grabowsky, Tara • Payne, Ron
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
A method, computer program product, and system continually obtain machine-readable data sets related to a patient population diagnosed with a medical condition from one or more databases (different computing nodes in the distributed environment). The processor(s) continually applies a recurrent neural network to the plurality of data sets to machine learn optimal features for classifying patients into multiple categories related to presence or progression of the medical condition. The processor(s) continually generates, based on the machine learned optimal set of features, intermediate features, based on the weightings of a portion of the machine learned optimal set of features, i.e., a model. The processor(s) evaluates the records and classifies records into the categories, based on a current model (model generated in real-time).
Core Innovation
The invention is a computer-implemented, distributed machine-learning framework for an orphan disease that continually obtains electronic medical records comprising a plurality of machine-readable data sets related to a patient population diagnosed with the medical condition from one or more databases. The machine-readable data comprising the plurality of machine-readable data sets are obtained from different computing nodes in the distributed computing environment. The method continually applies a neural network to the plurality of machine-readable data sets to machine learn an optimal set of features for classifying patients into a plurality of categories related to presence or progression of the medical condition.
The method continually generates intermediate features based on the weightings of a portion of the machine learned optimal set of features, where the intermediate features comprise a model of the medical condition. The method then obtains, at a given time, one or more data sets related to a patient population not diagnosed with the medical condition and evaluates a portion of records and classifies the portion of the records into the plurality of categories related to the medical condition. Each category represents a likelihood of having or developing the medical condition during a defined timeline based on a current model, and the current model is a version of the model generated in real-time based on the given time.
At a second given time, the method obtains second one or more data sets related to patients not diagnosed with the medical condition and classifies a portion of records based on a new current model that is a version of the model generated in real-time at the second given time. The new current model is different from the current model based on changes in machine-readable data between the given time and the second given time, and the intermediate features automatically change temporally based on those changes over time. Based on a frequency of features, the method identifies additional common features and weights the additional common features based on frequency of occurrence, where the additional common features comprise mutual information with values above a predefined threshold, selects a smallest subset of features that collectively contain a majority of the mutual information, generates patterns from the subset, and utilizes support vector machines and one or more classifier algorithms to identify presence or absence of the medical condition in an undiagnosed patient using features comprising the patterns.
Claims Coverage
The partial content provides three independent claim sets: a computer-implemented method, a computer program product, and a system. The claims describe the same core inventive construct of continually updating a neural-network-based model in a distributed computing environment to classify undiagnosed patients of an orphan disease into multiple likelihood/progression categories over a defined timeline, including mutual-information feature identification and pattern-based use of support vector machines and classifier algorithms.
Distributed continual obtaining of electronic medical records for an orphan disease
Continually obtaining electronic medical records comprising a plurality of machine-readable data sets related to a patient population diagnosed with an orphan disease, wherein machine-readable data comprising the plurality of machine-readable data sets are obtained from different computing nodes in the distributed computing environment, from one or more databases.
Neural network feature learning with weighted optimal feature set
Continually applying a neural network to the plurality of machine-readable data sets to machine learn an optimal set of features for classifying patients into a plurality of categories related to presence or progression of the orphan disease, wherein the machine learned optimal set of features comprise features identified by the neural network as occurring over the plurality of machine-readable data sets and weighted by the neural network.
Intermediate features as a temporally updated model generated in real-time
Continually generating, based on the machine learned optimal set of features, intermediate features based on the weightings of a portion of the machine learned optimal set of features, wherein the intermediate features comprise a model of the orphan disease; obtaining data sets at a given time for a patient population not diagnosed with the medical condition and classifying records into categories representing a likelihood during a defined timeline based on a current model generated in real-time at the given time; obtaining second data sets at a second given time and classifying into categories based on a new current model generated in real-time, wherein the new current model is different based on changes in machine-readable data and the intermediate features automatically change temporally based on those changes.
Mutual-information feature selection and pattern generation
Based on a frequency of features in the plurality of machine-readable data sets, identifying additional common features and weighting the additional common features based on frequency of occurrence, wherein the additional common features comprise mutual information with mutual information values above a predefined threshold; selecting a portion of the additional common features comprising a smallest subset of features from the one or more features that collectively contain a majority of the mutual information; generating one or more patterns comprising the portion of the additional common features.
Support vector machines and classifier algorithms tuned to the current model
Generating, utilizing one or more support vector machines and one or more classifier algorithms based on the one or more patterns comprising the portion of the additional common features, the one or more classifier algorithms to identify presence or absence of the medical condition in an undiagnosed patient based on absence or presence of features comprising the one or more patterns comprising the portion of the additional common features in data related to the undiagnosed patient; tuning, based on the current model, the one or more classifier algorithms; obtaining a third one or more data sets and classifying records into the plurality of categories related to the medical condition based on the one or more tuned classifier algorithms.
Across the independent claims, the inventive scope centers on continual distributed ingestion of electronic medical records for an orphan disease, neural-network-driven learning of a weighted optimal feature set, continual generation of intermediate features comprising a real-time-updated model that changes temporally with data, mutual-information-based feature selection into a smallest subset, pattern generation, and the use of support vector machines and classifier algorithms tuned to the current model to classify undiagnosed patients into multiple likelihood/progression categories over a defined timeline.
Stated Advantages
Earlier symptom/predictor identification.
Improved computational efficiency via distributed query distribution and information-theoretic feature selection.
Documented Applications
Classifying undiagnosed patients related to an orphan disease into a plurality of categories representing likelihood of having or developing the medical condition during a defined timeline, with categories related to presence or progression and with real-time model updates at given times.
Interested in licensing this patent?