Prompt-based language models for generating multi-modal electronic records

Inventors

Sun, Jimeng • Wang, Zifeng

Assignees

University of Illinois System

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-12541620-B2

Patent

Publication Date

2026-02-03

Expiration Date


Abstract

An example embodiment may involve obtaining text-based, ground truth electronic health records (EHRs), wherein the ground truth EHRs specify a sequence of medical visits involving a plurality of modalities, and wherein each of the medical visits specifies tokens representing at least one of the modalities; generating a training data set by perturbing the ground truth EHRs, wherein perturbing the ground truth EHRs involves deleting or shuffling some of the tokens in the ground truth EHRs; and iteratively applying a machine learning trainer application to the training data set, wherein the machine learning trainer application includes: (i) a bidirectional language model encoder that takes EHRs within the training data set and produces vector embeddings therefrom, (ii) an autoregressive language model decoder that takes the vector embeddings and infers predicted EHRs therefrom, and (iii) a loss function that compares the predicted EHRs to their corresponding ground truth EHRs.

Core Innovation

The invention obtains text-based, ground truth electronic health records (EHRs) that specify a sequence of medical visits involving a plurality of modalities, where each visit specifies tokens representing at least one of the modalities. The ground truth EHRs are perturbed to generate a training data set by deleting or shuffling some of the tokens in the ground truth EHRs. The perturbed token sequences support training that targets prediction of EHR content from corrupted inputs.

The invention iteratively applies a machine learning trainer application to the training data set. The trainer application includes a bidirectional language model encoder that takes the EHRs within the training data set and produces vector embeddings, and an autoregressive language model decoder that takes the vector embeddings and infers predicted EHRs therefrom. A loss function compares the predicted EHRs to their corresponding ground truth EHRs, and an updating function updates the bidirectional language model encoder or the autoregressive language model decoder based on output of the loss function.

In response to completion of the machine learning trainer application, the bidirectional language model encoder and the autoregressive language model decoder are provided as a generative model that can produce synthetic EHRs based on input EHRs provided thereto. The invention then produces, by way of the generative model, a privacy-preserving synthetic EHR and provides the privacy-preserving synthetic EHR to a downstream healthcare application.

Claims Coverage

The independent claims cover three implementations: a computer-implemented method, a non-transitory computer-readable medium storing program instructions, and a computing device. Across these, the inventive features are a perturbed tokenized, multi-modal EHR training data set, bidirectional encoder and autoregressive decoder training with embedding-to-prediction loss, and a generative model that produces privacy-preserving synthetic EHRs for a downstream healthcare application.

Perturbed tokenized, multi-modal EHR training data set

obtaining text-based, ground truth electronic health records (EHRs) that specify a sequence of medical visits involving a plurality of modalities, generating a training data set by perturbing the ground truth EHRs, wherein perturbing involves deleting or shuffling some of the tokens in the ground truth EHRs

Bidirectional encoder and autoregressive decoder with embedding-to-prediction loss training

iteratively applying a machine learning trainer application to the training data set, wherein the trainer application includes a bidirectional language model encoder that takes EHRs and produces vector embeddings, an autoregressive language model decoder that takes the vector embeddings and infers predicted EHRs, a loss function that compares the predicted EHRs to corresponding ground truth EHRs, and an updating function that updates the bidirectional language model encoder or the autoregressive language model decoder based on output of the loss function

Generative model that produces privacy-preserving synthetic EHRs for downstream healthcare application

in response to completion of the machine learning trainer application, providing the bidirectional language model encoder and the autoregressive language model decoder as a generative model that can produce synthetic EHRs based on input EHRs, producing by way of the generative model a privacy-preserving synthetic EHR, and providing the privacy-preserving synthetic EHR to a downstream healthcare application

Program instructions for the perturbation-to-loss generative training workflow

a non-transitory computer-readable medium having stored program instructions that, upon execution, cause operations comprising obtaining the text-based ground truth EHRs, generating a training data set by deleting or shuffling tokens, iteratively applying a machine learning trainer application with a bidirectional language model encoder, an autoregressive language model decoder, a loss function comparing predicted to ground truth EHRs, and an updating function updating encoder or decoder, providing the encoder and decoder as a generative model, producing a privacy-preserving synthetic EHR, and providing the privacy-preserving synthetic EHR to a downstream healthcare application

Computing device with components to execute the perturbation-to-loss generative training workflow

a computing device comprising one or more processors and memory, and program instructions that cause operations comprising obtaining text-based ground truth EHRs, generating a training data set by deleting or shuffling tokens, iteratively applying a machine learning trainer application with a bidirectional language model encoder, an autoregressive language model decoder, a loss function comparing predicted to ground truth EHRs, and an updating function updating encoder or decoder, providing the encoder and decoder as a generative model, producing a privacy-preserving synthetic EHR, and providing the privacy-preserving synthetic EHR to a downstream healthcare application

All independent claims share the same core claim coverage: training a generative model by perturbing token sequences from text-based, multi-modal ground truth EHRs, using a bidirectional language model encoder to produce vector embeddings and an autoregressive language model decoder to infer predicted EHRs, optimizing with a loss comparing predicted and ground truth EHRs via an updating function, and then generating privacy-preserving synthetic EHRs for downstream healthcare applications.

Stated Advantages

Not explicitly described in patent.

Documented Applications

Providing the privacy-preserving synthetic EHR to a downstream healthcare application.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.