Pipelined operations in neural networks
Inventors
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
An integrated circuit (IC) implements an M by N aperture function over an R by C source array. The IC has an input port receiving an ordered stream of independent input values, an output port producing an output stream, a mass multiplier circuit multiplying inputs by weights, producing streams of products on pathways on the IC, an M by N array of compositor circuits on the IC, single dedicated pathways between compositors, delay circuits, a finalization circuit, and a control circuit operating counters and producing control signals. The compositors combine the values received from product pathways, further combine that result to an initial value or to a value from an adjacent compositor upstream, or to a value from a delay circuit. Upon a last downstream compositor producing a complete composition of values, that value is passed to the finalization circuit, which posts a result to the output port.
Core Innovation
The invention relates to an integrated circuit implementing an M by N aperture function over an R by C source array to produce an R by C destination array. The IC receives an ordered stream of independent input values from the source array through an input port, and produces an ordered output stream of output values into the destination array through an output port. A mass multiplier circuit multiplies, in parallel and in order, each input value by every weight required by the aperture function, producing streams of products on parallel conductive product pathways, where each product pathway is dedicated to a single product of an input by a weight value.
The IC includes an M by N array of compositor circuits, where each compositor is associated with a sub-function of the aperture function at an (m,n) position and is coupled by a dedicated pathway to each product pathway carrying a product produced from a weight value associated with the sub-function. Delay circuits receive values on dedicated pathways from compositors and provide those values delayed at later times on dedicated pathways to other compositors downstream. A finalization circuit is provided, and a control circuit operating counters produces control signals coupled to the compositors, the delay circuits, and the finalization circuit.
For each source interval, the compositors combine the values received from the dedicated connections to the parallel conductive pathways, further combine that result to an initial value for that compositor or to a value on the dedicated pathway from an adjacent compositor upstream, or to a value received from a delay circuit, and post the combined result to a register coupled to the dedicated pathway to the adjacent compositor downstream, or to a delay circuit, or both. Upon a last downstream compositor producing a complete composition of values for an output of the aperture function at a specific position of the R by C array of inputs, that composed value is passed to the finalization circuit, which processes the value and posts the result to the output port as one value of the output stream.
Claims Coverage
The partial content includes two independent claims: an apparatus claim for an integrated circuit, and a method claim for implementing the M by N aperture function over an R by C source array. Across these independent claims, there are core inventive features related to parallel mass multiplication, an M by N compositor array with dedicated pathways, delay circuits for time-aligned downstream compositing, and a finalization circuit producing ordered output stream values under control signals from a control circuit executing counters.
Streaming input and ordered output for an M by N aperture function
An IC implementing an M by N aperture function over an R by C source array to produce an R by C destination array, with an input port receiving an ordered stream of independent input values from the source array and an output port producing an ordered output stream of output values into the destination array.
Parallel mass multiplication into dedicated product pathways
A mass multiplier circuit coupled to the input port, multiplying in parallel each input value in order by every weight required by the aperture function, producing streams of products on a set of parallel conductive product pathways on the IC, each product pathway dedicated to a single product of an input by a weight value.
M by N compositor array with dedicated pathways to product streams
An M by N array of compositor circuits on the IC, each compositor circuit associated with a sub-function of the aperture function at the (m,n) position, and coupled by a dedicated pathway to each of the set of product pathways carrying a product produced from a weight value associated with the sub-function, with single dedicated pathways between compositors.
Delay circuits for downstream compositing and value alignment
Delay circuits on the IC receiving values on dedicated pathways from compositors and providing the values delayed at later times on dedicated pathways to other compositors downstream.
Control circuit executing counters for compositors, delays, and finalization
A control circuit operating counters and producing control signals coupled to the compositors, the delay circuits, and the finalization circuit, with compositors combining values received from dedicated connections and further combining that result to an initial value, an adjacent compositor upstream value, or a value received from a delay circuit.
Finalization based on last downstream compositor producing complete composition
Upon a last downstream compositor producing a complete composition of values for an output of the aperture function at a specific position of the R by C array of inputs, that composed value is passed to the finalization circuit, which processes the value and posts the result to the output port as one value of the output stream.
Mass-multiplier parallel product streaming during method execution
A method implementing an M by N aperture function over an R by C source array, producing an R by C destination array, comprising providing an ordered stream of independent input values from the source array to an input port of an integrated circuit (IC) and multiplying in parallel each input value in order by every weight value required by the aperture function by a mass multiplier circuit on the IC coupled to the input port, producing streams of products on a set of parallel conductive product pathways, each product pathway dedicated to a single product of an input by a weight value.
Compositor array with dedicated connections, delay circuits, and finalization
Providing to each of an M by N array of compositor circuits on the IC, each compositor circuit associated with a sub-function of the aperture function, by dedicated connections to each compositor circuit from the streams of products, those products produced from a weight value associated with the sub-function, providing control signals to the compositors, a plurality of delay circuits and a finalization circuit by a control circuit executing counters, combining in each source cycle the values received from the dedicated connections with an initial value or an upstream dedicated-pathway value or a value received from one of the plurality of delay circuits, and upon a last downstream compositor producing a complete combination of values for an output at a specific position, providing that complete combination to a finalization circuit.
Ordered output stream produced by finalization while continuing until last output value
Processing the complete combination by the finalization circuit and posting the result to an output port as one value in an ordered output stream, and continuing operation of the IC until all input elements have been received and a last output value has been produced to the output stream.
Across the independent claims, the apparatus and method are centered on an IC that streams ordered inputs and outputs, performs parallel mass multiplication of each input by every required weight into dedicated product pathways, performs sub-function compositing using an M by N compositor array with dedicated single-pathway connections, aligns values using delay circuits, and produces each ordered output stream value via a finalization circuit when a last downstream compositor completes a complete composition for a specific output position.
Stated Advantages
Avoids RAM buffering and reads by using single-pass streaming with stored partial sums [described at a conceptual level in the provided partial content].
Enables high throughput and pixel-synchronous pipeline operation [described at a conceptual level in the provided partial content].
Supports integration of truncated edge results (left/right and top/bottom regions) into the streaming computation [as stated in the provided partial content].
Allows omission of specific outputs from the output stream in a fixed or variable stepping pattern [as stated in the provided partial content].
Documented Applications
Hardware IC implementation of an M by N aperture function over R by C arrays, including convolutional neural node behavior and streaming output generation [as stated in the provided partial content].
Hardware-assisted neural network training system aspects, including extension to 3D apertures and discussion of gradient/precision via cascaded mass multipliers [as stated in the provided partial content].
Interested in licensing this patent?