Implementation of machine-learning based query construction and pattern identification through visualization in user interfaces

Inventors

Miller, ChrisFolta, TylerGrabowsky, TaraShukla, Oodaye

Assignees

Eversana Life Science Services LLCLscs Holdings Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-11450434-B2

Patent

Publication Date

2022-09-20

Expiration Date


Abstract

A computer system, computer-implemented method, and computer program product include a processor(s) (executing code) that obtains a data set(s) related to a patient population diagnosed with a medical condition a database(s). The processor(s) identifies common features, generates patterns of the common features, and generates machine learning algorithms based on the patterns to identify presence or absence of the given medical condition in an undiagnosed patient. The processor(s) compiles a training set of data and tunes the machine learning algorithms with the training set of data. The processor(s) integrates the machine learning algorithms into a graphical user interface. The processor(s) obtains data related to the undiagnosed patient via the interface and applies the machine learning algorithms to determine a probability (numerical value indicating a percentage of commonality between the data related to the undiagnosed patient and the one or more patterns) and display the probability as a score in the interface.

Core Innovation

A distributed computing environment obtains one or more data sets related to a patient population diagnosed with a medical condition from one or more databases. Based on a frequency of features in the one or more data sets, the one or more processors identify common features in the one or more data sets and weight the common features based on frequency of occurrence in the portion of the data, wherein the common features comprise mutual information. The mutual information is used to generate a patient definition by truncating the common features using mutual information values above a predefined threshold.

The patient definition is formed by selecting, from the common features with mutual information values above the predefined threshold, a portion of the common features that comprises a smallest subset of common features with the mutual information values above the predefined threshold comprising a majority of common features of the one or more features with the mutual information values above the predefined threshold. One or more machine learning algorithms are generated based on the patient definition to identify presence or absence of the given medical condition in an undiagnosed patient.

The machine learning algorithms are tuned by applying the one or more machine learning algorithms to a training set of data. The training set comprises data from the one or more data sets and at least one additional data set comprising data related to a population without the medical condition, and the machine learning algorithms are integrated into a graphical user interface that provides an input for a user to provide data related to the undiagnosed patient.

After the data related to the undiagnosed patient is provided, the machine learning algorithms are applied to determine a probability. The probability is a numerical value indicating a percentage of commonality between the data related to the undiagnosed patient and the patient definition, and the probability is displayed to the user through the graphical user interface as a score.

Claims Coverage

The document provides three independent claims. Across these claims, there are three principal inventive features centered on mutual-information-weighted feature selection with threshold-based truncation into a patient definition, machine-learning training using statistical sampling that includes a population without the medical condition, and integration into a graphical user interface to produce and display a probability/score for an undiagnosed patient.

Mutual information-weighted common features to generate a patient definition

Identifying common features in the one or more data sets and weighting the common features based on frequency of occurrence in the portion of the data, wherein the common features comprise mutual information; utilizing the mutual information to generate a patient definition by truncating the common features based on identifying one or more common features with mutual information values above a predefined threshold; and selecting a portion of the common features comprising a smallest subset of common features with the mutual information values above the predefined threshold comprising a majority of common features.

Machine learning algorithms trained using statistical sampling with a population without the medical condition

Generating one or more machine learning algorithms based on the patient definition to identify presence or absence of the given medical condition in an undiagnosed patient; utilizing statistical sampling to compile a training set of data wherein the training set comprises data from the one or more data sets and at least one additional data set comprising data related to a population without the medical condition; and tuning the one or more machine learning algorithms by applying the one or more machine learning algorithms to the training set of data.

Graphical user interface integrated scoring for probability based on percentage commonality

Integrating the one or more machine learning algorithms into a graphical user interface that provides an input to enable a user to provide data related to the undiagnosed patient; applying the one or more machine learning algorithms to the data related to the undiagnosed patient; determining a probability indicating a percentage of commonality between the data related to the undiagnosed patient and the patient definition; and displaying the probability to the user through the graphical user interface as a score.

Across the independent claims, coverage centers on forming a patient definition from mutual information-weighted common features with predefined-threshold truncation, training machine learning algorithms using statistical sampling that includes a population without the medical condition, and presenting a probability/score through a graphical user interface for an undiagnosed patient.

Stated Advantages

Not explicitly described in patent.

Documented Applications

Not explicitly described in patent.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.