Prompt-based language models for generating multi-modal electronic records
Inventors
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
An example embodiment may involve obtaining text-based, ground truth electronic health records (EHRs), wherein the ground truth EHRs specify a sequence of medical visits involving a plurality of modalities, and wherein each of the medical visits specifies tokens representing at least one of the modalities; generating a training data set by perturbing the ground truth EHRs, wherein perturbing the ground truth EHRs involves deleting or shuffling some of the tokens in the ground truth EHRs; and iteratively applying a machine learning trainer application to the training data set, wherein the machine learning trainer application includes: (i) a bidirectional language model encoder that takes EHRs within the training data set and produces vector embeddings therefrom, (ii) an autoregressive language model decoder that takes the vector embeddings and infers predicted EHRs therefrom, and (iii) a loss function that compares the predicted EHRs to their corresponding ground truth EHRs.
Core Innovation
The invention obtains text-based, ground truth electronic health records (EHRs) that specify a sequence of medical visits involving a plurality of modalities, where each visit specifies tokens representing at least one of the modalities. The ground truth EHRs are perturbed to generate a training data set by deleting or shuffling some of the tokens in the ground truth EHRs. The perturbed token sequences support training that targets prediction of EHR content from corrupted inputs.
The invention iteratively applies a machine learning trainer application to the training data set. The trainer application includes a bidirectional language model encoder that takes the EHRs within the training data set and produces vector embeddings, and an autoregressive language model decoder that takes the vector embeddings and infers predicted EHRs therefrom. A loss function compares the predicted EHRs to their corresponding ground truth EHRs, and an updating function updates the bidirectional language model encoder or the autoregressive language model decoder based on output of the loss function.
In response to completion of the machine learning trainer application, the bidirectional language model encoder and the autoregressive language model decoder are provided as a generative model that can produce synthetic EHRs based on input EHRs provided thereto. The invention then produces, by way of the generative model, a privacy-preserving synthetic EHR and provides the privacy-preserving synthetic EHR to a downstream healthcare application.
Claims Coverage
The independent claims cover three implementations: a computer-implemented method, a non-transitory computer-readable medium storing program instructions, and a computing device. Across these, the inventive features are a perturbed tokenized, multi-modal EHR training data set, bidirectional encoder and autoregressive decoder training with embedding-to-prediction loss, and a generative model that produces privacy-preserving synthetic EHRs for a downstream healthcare application.
Perturbed tokenized, multi-modal EHR training data set
obtaining text-based, ground truth electronic health records (EHRs) that specify a sequence of medical visits involving a plurality of modalities, generating a training data set by perturbing the ground truth EHRs, wherein perturbing involves deleting or shuffling some of the tokens in the ground truth EHRs
Bidirectional encoder and autoregressive decoder with embedding-to-prediction loss training
iteratively applying a machine learning trainer application to the training data set, wherein the trainer application includes a bidirectional language model encoder that takes EHRs and produces vector embeddings, an autoregressive language model decoder that takes the vector embeddings and infers predicted EHRs, a loss function that compares the predicted EHRs to corresponding ground truth EHRs, and an updating function that updates the bidirectional language model encoder or the autoregressive language model decoder based on output of the loss function
Generative model that produces privacy-preserving synthetic EHRs for downstream healthcare application
in response to completion of the machine learning trainer application, providing the bidirectional language model encoder and the autoregressive language model decoder as a generative model that can produce synthetic EHRs based on input EHRs, producing by way of the generative model a privacy-preserving synthetic EHR, and providing the privacy-preserving synthetic EHR to a downstream healthcare application
Program instructions for the perturbation-to-loss generative training workflow
a non-transitory computer-readable medium having stored program instructions that, upon execution, cause operations comprising obtaining the text-based ground truth EHRs, generating a training data set by deleting or shuffling tokens, iteratively applying a machine learning trainer application with a bidirectional language model encoder, an autoregressive language model decoder, a loss function comparing predicted to ground truth EHRs, and an updating function updating encoder or decoder, providing the encoder and decoder as a generative model, producing a privacy-preserving synthetic EHR, and providing the privacy-preserving synthetic EHR to a downstream healthcare application
Computing device with components to execute the perturbation-to-loss generative training workflow
a computing device comprising one or more processors and memory, and program instructions that cause operations comprising obtaining text-based ground truth EHRs, generating a training data set by deleting or shuffling tokens, iteratively applying a machine learning trainer application with a bidirectional language model encoder, an autoregressive language model decoder, a loss function comparing predicted to ground truth EHRs, and an updating function updating encoder or decoder, providing the encoder and decoder as a generative model, producing a privacy-preserving synthetic EHR, and providing the privacy-preserving synthetic EHR to a downstream healthcare application
All independent claims share the same core claim coverage: training a generative model by perturbing token sequences from text-based, multi-modal ground truth EHRs, using a bidirectional language model encoder to produce vector embeddings and an autoregressive language model decoder to infer predicted EHRs, optimizing with a loss comparing predicted and ground truth EHRs via an updating function, and then generating privacy-preserving synthetic EHRs for downstream healthcare applications.
Stated Advantages
Not explicitly described in patent.
Documented Applications
Providing the privacy-preserving synthetic EHR to a downstream healthcare application.
Interested in licensing this patent?