System and method for training neural networks

Inventors

XIONG, Hui Yuan • Delong, Andrew • Frey, Brendan

Assignees

University of Toronto • Deep Genomics Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-11681917-B2

Patent

Publication Date

2023-06-20

Expiration Date


Abstract

Systems and methods for training a neural network or an ensemble of neural networks are described. A hyper-parameter that controls the variance of the ensemble predictors is used to address overfitting. For larger values of the hyper-parameter, the predictions from the ensemble have more variance, so there is less overfitting. This technique can be applied to ensemble learning with various cost functions, structures and parameter sharing. A cost function is provided and a set of techniques for learning are described.

Core Innovation

The invention provides a variance-adjustable dropout and ensemble training approach in which an aggregate output is generated from a plurality of outputs produced by applying a plurality of neural networks to a training data item. For each neural network output, the method computes a difference between the network output and the aggregate output, scales the difference by a hyper-parameter, and adds the scaled difference to the aggregate output to generate a variance-adjusted output. The method then computes a difference between the variance-adjusted output and a pre-determined output for the training data item, and configures the plurality of neural networks to reduce that difference.

A variance-adjusted predictor is constructed so that the output is expressed using a mean prediction and a dropout-induced prediction term whose contribution is modulated by the hyper-parameter. During training, gradients combine contributions from the mean-network portion and the dropout-network portion, where the hyper-parameter controls the relative variance of dropout-induced predictions around the mean. The hyper-parameter is selected using cross-validation, including criteria such as minimizing validation error or maximizing validation log-likelihood.

The approach implements dropout and ensemble behavior by disabling hidden units and/or input units of the neural network using random dropout probabilities, pseudo-random dropout probabilities, predetermined sets of binary masks used only once, or fixed patterns with predetermined probabilities. An ensemble is formed by repeatedly applying the neural network with the disabled or masked units and aggregating the resulting outputs, with aggregation at test time described as averaging or majority voting.

Claims Coverage

The independent claims are clm-00001 and clm-00021. They each center on producing an aggregate output, computing a variance-adjusted output using a hyper-parameter, and configuring the neural network(s) to reduce a difference relative to a pre-determined output; clm-00021 additionally specifies unit disabling during repeated application using random, pseudo-random, one-time masks, or fixed patterns.

Variance-adjusted training outputs using a hyper-parameter relative to an aggregate output

computing a difference between an output of the neural network and the aggregate output, scaling the difference by a hyper-parameter, adding the scaled difference to the aggregate output to generate a variance-adjusted output, and computing a difference between the variance-adjusted output and a pre-determined output for the training data item.

Configuring multiple neural networks to reduce the variance-adjusted output difference

configuring the plurality of neural networks to reduce the difference between the variance-adjusted output and the pre-determined output.

Aggregate output generated from multiple neural network outputs applied to a training data item

using a plurality of outputs to generate an aggregate output, wherein the plurality of outputs are generated at least in part by applying the plurality of neural networks to a training data item.

Repeating application while disabling hidden unit or input unit using random, pseudo-random, one-time binary masks, or fixed pattern

applying the neural network to the training data item comprises disabling at least one hidden unit or input unit of the neural network randomly, pseudo-randomly, using a predetermined set of binary masks used only once, or according to a fixed pattern, each with a predetermined probability.

Variance-adjusted training outputs for each plurality of training outputs relative to a predetermined training output

for each of the plurality of training outputs, computing a difference between the training output and the aggregate output, scaling the difference by a hyper-parameter, adding the scaled difference to the aggregate training output to generate a variance-adjusted training output, and computing a difference between the variance-adjusted training output and a pre-determined training output for the training data item.

Configuring the neural network to reduce variance-adjusted training output differences

configuring the neural network to reduce the difference between the variance-adjusted training outputs and the pre-determined training output.

Across clm-00001 and clm-00021, the claims cover generating an aggregate output from multiple neural network outputs, computing variance-adjusted outputs by scaling the difference between each output and the aggregate output with a hyper-parameter, and configuring the neural network(s) to reduce the difference between the variance-adjusted output(s) and a pre-determined output; clm-00021 further ties repeated application to disabling hidden or input units via random, pseudo-random, one-time binary masks, or fixed patterns.

Stated Advantages

Provides training and test aggregation using averaging or majority voting from repeated applications with unit disabling or masking.

Improves held-out performance indicators described in the partial content, where MNIST results suggest that a selected hyper-parameter value improves held-out error.

Documented Applications

Digit classification on MNIST, with results described for held-out error under a selected hyper-parameter.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.