Systems and methods for deep model translation generation

Inventors

Kaufhold, John PatrickSleeman, Jennifer Alexander

Assignees

General Dynamics Mission Systems Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-10504004-B2

Patent

Publication Date

2019-12-10

Expiration Date


Abstract

Embodiments of the present invention relate to systems and methods for improving the training of machine learning systems to recognize certain objects within a given image by supplementing an existing sparse set of real-world training images with a comparatively dense set of realistic training images. Embodiments may create such a dense set of realistic training images by training a machine learning translator with a convolutional autoencoder to translate a dense set of synthetic images of an object into more realistic training images. Embodiments may also create a dense set of realistic training images by training a generative adversarial network (“GAN”) to create realistic training images from a combination of the existing sparse set of real-world training images and either Gaussian noise, translated images, or synthetic images. The created dense set of realistic training images may then be used to more effectively train a machine learning object recognizer to recognize a target object in a newly presented digital image.

Core Innovation

The invention improves the training of a computer-based object recognizer to recognize an object within an image by using both real-world images and synthetic imagery created by two coupled components. A set of real-world images of a target object is obtained, and a set of synthetic images of the target object is created.

A translator is trained to produce translated images of the target object, and the translator comprises a first computer-based machine learning system having a convolutional autoencoder. The translator is trained using pairings obtained by identifying a real-world image that corresponds with a synthetic image, and the trained translator is invoked to produce a plurality of translated images of the target object.

In parallel, a GAN is trained with a second computer-based machine learning system comprising a discriminative neural network and a generative neural network. GAN training includes instantiating the generative neural network with Gaussian noise and instantiating the discriminative neural network with the set of real-world images, where each real-world image is labeled according to the target object.

The object recognizer is trained using a collection assembled from real-world images, the translated images, and the generated images to recognize the target object within a newly presented digital image obtained from an external image sensor. The disclosed system also includes pairing strategies and coupling translator loss with object-recognizer activations to synchronize training.

Claims Coverage

Independent claim clm-00001 provides the core end-to-end coverage, and dependent claims clm-00004, clm-00003, clm-00006, clm-00008, and clm-00009 narrow or extend the pipeline by constraining data sources used for translator/GAN components and by adding synchronized loss usage. The claim set identifies 8 inventive features.

Real-world-to-synthetic translator with convolutional autoencoder pairings

A translator trained to produce a plurality of translated images of the target object, where the translator comprises a first computer-based machine learning system having a convolutional autoencoder, where translator training includes providing the translator with a plurality of pairings obtained by identifying one real-world image that corresponds with one synthetic image and learning from the pairings how to produce translated images of the target object.

GAN training with Gaussian-noise instantiation and labeled real-world discrimination

A generative adversarial network (GAN) trained to produce a plurality of generated images of the target object, where the GAN comprises a second computer-based machine learning system having a discriminative neural network and a generative neural network, where GAN training includes instantiating the generative neural network with Gaussian noise and instantiating the discriminative neural network with the set of real-world images labeled according to the target object.

Object recognizer training on combined real-world, translated, and generated image collections for external sensing

Training an object recognizer to recognize the target object within a newly presented digital image by providing the object recognizer a collection of training images assembled from the set of real-world images, the plurality of translated images, and the plurality of generated images, where the object recognizer is embodied in at least a portion of a high-speed graphics processing unit of a digital computer, and using the trained object recognizer to recognize the target object in the new digital image obtained from an external image sensor.

Second pairing training that matches real-world images to generated images

Training a translator by providing multiple second pairings formed by matching real-world images to generated images and using these second pairings to train a convolutional autoencoder to produce a plurality of translated images of the target object.

GAN discriminator instantiation with translated images

GAN training includes instantiating the discriminative neural network using a set of translated images.

Translator and GAN constraints using generated and synthetic image sources

Training a translator by providing it with generated images and training a GAN by instantiating the discriminative neural network with a set of synthetic images.

Synthetic images derived from generated images

Obtaining a set of synthetic images of the target object from generated images.

Synchronized translator and object-recognizer training using first and second loss functions

Synchronized training between a translator using a convolutional autoencoder with an iteratively invoked first loss function and an object recognizer with an iteratively invoked second loss function, where second-loss values are used in computations for the first loss function.

The claim set is centered on improving object-recognizer training by combining real-world images with translated images produced by a convolutional autoencoder translator and generated images produced by a GAN, then training the object recognizer on a merged collection for use on newly sensed images. Dependent claims further specify alternative pairing/data-source constraints for translator and GAN components and add synchronized training via cross-used loss values.

Stated Advantages

Reduced real-world data acquisition.

Improved accuracy with fewer labeled samples.

Filling in sparse or missing classes.

Documented Applications

Recognizing a target object within newly presented digital images obtained from an external image sensor by using a trained object recognizer.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.