Image generation with flexible sampler for diffusion modeling

Inventors

Wang, WujieIngraham, John

Assignees

Generate Biomedicines Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-12579706-B2

Patent

Publication Date

2026-03-17

Expiration Date


Abstract

A system performs image generation with a flexible sampler used in a diffusion model. The system receives a set of sampling parameter(s) to tailor sampling. The sampling parameter(s) include a combination of a momentum sampling parameter, a mass sampling parameter, a friction sampling parameter, an equilibration sampling parameter, and an inverse temperature sampling parameter. The system modifies a sampling operation based on the sampling parameter(s). The system applies a diffusion model with the modified sampling operation to generate a synthetic image. The system may present the synthetic image on a graphical user interface.

Core Innovation

The invention provides a computer-implemented method for generating a synthetic image based on a text prompt by using a diffusion model whose sampling operation is modified by a set of sampling parameters. The method receives the text prompt via a graphical user interface, generates one or more embeddings for the text prompt in an embedding space, and receives sampling parameters that tailor sampling, including momentum, mass, friction, equilibration, and inverse temperature.

The method modifies a sampling operation based on the set of sampling parameters to yield a first modified sampling operation. In each sampling step of a plurality of sampling steps, the diffusion model denoises a prior sampled state based on the one or more embeddings of the text prompt, and then samples a subsequent sampled state from the denoised prior sampled state by applying the first modified sampling operation. The final sampled state is the synthetic image that satisfies the features of the text prompt.

The same flexible sampling framework is applied to structured data objects generated from a text prompt. The method receives a text prompt via a graphical user interface, generates one or more embeddings in an embedding space, receives the set of sampling parameters, and modifies a sampling operation to a first modified sampling operation. The diffusion model denoises a prior sampled state of the structured data object based on the embeddings and samples subsequent sampled states using the first modified sampling operation until a final sampled state satisfies the prompt features.

Claims Coverage

The partial content includes three independent claims. Across these independent claims, the inventive coverage centers on receiving a GUI text prompt and generating prompt embeddings, receiving and combining momentum, mass, friction, equilibration, and inverse temperature sampling parameters to modify a sampling operation, and applying a diffusion model that, at each sampling step, denoises a prior sampled state based on the embeddings and then samples a subsequent sampled state using the modified sampling operation.

Gui text prompt with embedding-space representation

Receiving, via a graphical user interface, a text prompt that specifies features of a synthetic image or a structured data object to be generated; generating one or more embeddings for the text prompt in an embedding space.

Sampling parameters combining momentum, mass, friction, equilibration, and inverse temperature

Receiving a set of one or more sampling parameters that tailor sampling, wherein the set includes a combination of a momentum sampling parameter, a mass sampling parameter, a friction sampling parameter, an equilibration sampling parameter, and an inverse temperature sampling parameter; modifying a sampling operation based on the set of sampling parameters to yield a first modified sampling operation.

Diffusion-model sampling with denoising of a prior sampled state and subsequent sampling using the modified operation

Applying a diffusion model to generate the synthetic image or structured data object, wherein in each sampling step of a plurality of sampling steps: denoising a prior sampled state based on the one or more embeddings of the text prompt; and sampling a subsequent sampled state from the denoised prior sampled state by applying the first modified sampling operation.

Final sampled state satisfies prompt features and gui presentation

A final sampled state is the synthetic image that satisfies the features of the text prompt, or the structured data object that satisfies the features of the text prompt; the synthetic image is provided for presentation via the graphical user interface, or a visual representation of the structured data object is rendered and provided for presentation.

Together, the independent claims cover a diffusion-model sampling workflow in which prompt embeddings guide denoising at each sampling step, while a modified sampling operation is tailored using a combined set of momentum, mass, friction, equilibration, and inverse temperature sampling parameters. The workflow produces a final sampled state that satisfies the text prompt features and presents a synthetic image or a rendered visual representation of a structured data object via the graphical user interface.

Stated Advantages

Not explicitly described in patent.

Documented Applications

Generating a synthetic image based on a text prompt specifying features of the synthetic image; providing the synthetic image for presentation via the graphical user interface.

Generating a structured data object based on a text prompt specifying features of the structured data object; rendering a visual representation of the structured data object; providing the visual representation for presentation via the graphical user interface.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.