Composing arbitrary convolutional neural network models from a fixed set of duplicate pipelined components

Inventors

Mathews, Mark Ashley

Assignees

Gigantor Technologies Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-11797832-B2

Patent

Publication Date

2023-10-24

Expiration Date


Abstract

An Application Specific Integrated Circuit (ASIC) for computing a convolutional neural network (CNN) has a first input bus receiving an ordered stream of values from an array, each position in the array having one or more channels, and a plurality of kernel processing tiles receiving inputs through configurable multiplexors. The kernel processing tiles and buses are arranged and connected in a manner that the ASIC operates as a pipelined system delivering an output stream in synchronization with the input stream.

Core Innovation

The invention provides an Application Specific Integrated Circuit (ASIC) for computing a convolutional neural network (CNN) convolution. The ASIC includes an input bus receiving an ordered stream of values from an array, where each position in the array has one or more data channels. The ASIC processes the ordered stream using multiple ordered sets of kernel processing tiles with fixed numbers of parallel input connections and parallel output connections.

The architecture is organized into a first ordered set, a second ordered set, and a third ordered set of kernel processing tiles arranged first to last. Each ordered set is coupled to a corresponding input bus stage through configurable multiplexors, and each tile passes computed values both to an output bus and to an adjacent downstream kernel processing tile of the same ordered set. The output bus is connected back as an input to each configurable multiplexor of that ordered set, enabling continued streaming and tile-to-tile propagation while maintaining fixed parallel interfaces.

A primary output path is formed by routing computed values from the third ordered set through a single primary output multiplexor to a primary output circuit adapted to perform primary output processing and provide final output. The architecture constrains the kernel processing to a common kernel size for the tiles in each ordered set and supports convolution computation for named common kernel sizes such as 3×3, 5×5, 7×7, or 9×9. The described implementation framework further includes activation via lookup table and auxiliary function tiles that support MaxPool, Average, Sample, and Expand, with interleaved multi-scale operation and N-up parallel processing.

Claims Coverage

Independent claim clm-00001 provides a streaming ASIC convolution architecture using three ordered sets of kernel processing tiles with configurable multiplexors, feedback through output buses, and a single primary output multiplexor feeding a primary output circuit. The dependent claims refine this architecture with specific common kernel sizes and additional auxiliary and routing circuitry, including pooling/processing tile function selection and specialized multiplexing for auxiliary outputs.

Ordered stream input bus with parallel data channels

An input bus receiving an ordered stream of values from an array, each position in the array having one or more data channels.

Three ordered sets of kernel processing tiles with configurable multiplexors and feedback buses

A first ordered set, first to last, of kernel processing tiles having a fixed number of parallel input connections and a fixed number of parallel output connections, each kernel processing tile of the first ordered set coupled to the input bus through one of a first set of configurable multiplexors, the kernel processing tiles adapted to compute a convolution for a common kernel size, and to pass the computed values both to a first output bus connected as an input back to each configurable multiplexor of the first set of configurable multiplexors and to an adjacent downstream kernel processing tile of the first ordered set; and a second ordered set, first to last, of kernel processing tiles coupled to the first output bus through one of a second set of configurable multiplexors, adapted to compute a convolution for the common kernel size, and to pass computed values both to a second output bus connected as an input back to each configurable multiplexor of the second set and to an adjacent downstream kernel processing tile of the second ordered set; and a third ordered set, first to last, coupled to the second output bus through one of a third set of configurable multiplexors, adapted to compute a convolution for the common kernel size, and to pass computed values both to a third output bus connected as an input back to each configurable multiplexor of the third set and to an adjacent downstream kernel processing tile of the third ordered set.

Primary output multiplexor and primary output circuit for final output

The third output bus connected though a single primary output multiplexor to a primary output circuit adapted to perform primary output processing and to provide final output.

Common kernel size constraints for selected convolution kernels

The ASIC configured such that the common kernel size is a 3×3 kernel and configured to compute 5×5, 7×7, or 9×9 convolutions.

Auxiliary function tile pooling/processing function selection

Each auxiliary function tile takes parallel inputs from two separate buses via a dual multiplexor and then outputs one selected pooling/processing function (MaxPool, Average, Sample, or Expand) chosen by an output multiplexer.

Specialized multiplexor routing for auxiliary output candidates

Input channel signals from separate parallel connections routed through a second specialized multiplexor that alternately samples and concatenates two parallel inputs into a single parallel output feed for selection by the output multiplexor as an auxiliary function tile output candidate.

Across the independent claim and its refinements, the ASIC centers on three sequential ordered sets of fixed-parallel kernel processing tiles that compute convolutions for a common kernel size and propagate computed values via adjacent tile chaining and feedback through configurable multiplexor input buses, culminating in a single primary output multiplexor and primary output circuit. Dependent refinements specify common kernel sizes and add auxiliary function tiles and specialized multiplexing/routing behaviors for pooling and auxiliary output candidates.

Stated Advantages

Throughput/latency improvements [procedural detail omitted for safety].

Reduced power/footprint.

Documented Applications

Hardware-assisted CNN convolution computation and primary output processing using a configurable ASIC convolution architecture.

Auxiliary function processing and activation via lookup table as part of CNN processing, including MaxPool, Average, Sample, and Expand.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.