Parallel implementation of deep neural networks applied to three-dimensional data sets

Inventors

Mathews, Mark Ashley

Assignees

Gigantor Technologies Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-11354571-B2

Patent

Publication Date

2022-06-07

Expiration Date


Abstract

An integrated circuit (IC) implements an L by M by N three-dimensional aperture function throughout a P by R by C three-dimensional source array. The IC has an input port receiving an ordered stream of independent input values from the source array, an output port producing an ordered stream of independent output values, an array of n compositor circuits, where n=L×M×N, each compositor circuit implementing a sub-function of the aperture function, dedicated pathways between the compositor circuits, delay circuits on the IC receiving values on the dedicated pathways from individual ones of the compositor circuits and providing the delayed values at later times to other compositor circuits downstream, a finalization circuit, and a control circuit operating counters and producing control signals coupled to the compositors, the delay circuits, and the finalization circuit.

Core Innovation

The invention implements an L by M by N three-dimensional aperture function throughout a P by R by C three-dimensional source array in an integrated circuit. The IC receives an ordered stream of independent input values from the source array via an input port, and produces an ordered stream of independent output values at an output port. The aperture function is decomposed into sub-functions using an array of n=L by M by N compositor circuits.

Dedicated pathways connect the compositor circuits, and delay circuits on the IC receive values on the dedicated pathways from individual ones of the compositor circuits. The delay circuits provide delayed values at later times to other compositor circuits downstream, supporting everted/inside-out order of operations across sub-functions of the aperture function. A finalization circuit sequences and completes the independent output values after processing compositions associated with specific positions in the P by R by C source array.

A control circuit operates counters and produces control signals that are coupled to the compositors, delay circuits, and finalization circuit. The disclosed approach includes streaming processing and ordered input/output behavior such that the output stream is produced while inputs are being received, without RAM buffering as described in the provided content. The architecture is also extended to three-dimensional convolutional computation examples using an everted aperture-function decomposition and streaming partial-sum retention concepts.

Claims Coverage

The independent claim defines an IC architecture implementing an L×M×N three-dimensional aperture function over a P×R×C three-dimensional source array using an ordered input stream/output stream and an array of n=L×M×N compositor circuits. The inventive features across dependent claims refine ordering, per-weight multiplication, finalization posting order, simultaneous streaming behavior, and quantitative constraints.

Integrated circuit implementing an L by M by N three-dimensional aperture function over a P by R by C three-dimensional source array

An IC implementing an L by M by N three-dimensional aperture function throughout a P by R by C three-dimensional source array using an input port receiving an ordered stream of independent input values, an output port producing an ordered stream of independent output values, and an array of n=L by M by N compositor circuits each implementing a sub-function of the aperture function.

Dedicated pathways between compositor circuits with downstream delay circuits

Dedicated pathways between the compositor circuits and delay circuits on the IC receiving values on the dedicated pathways from individual ones of the compositor circuits and providing the delayed values at later times to other compositor circuits downstream.

Finalization circuit for completing ordered output values

A finalization circuit in the IC.

Control circuit operating counters and producing control signals

A control circuit operating counters and producing control signals coupled to the compositors, the delay circuits, and the finalization circuit.

Mass multiplier circuit multiplying input-stream values by aperture weights

A mass multiplier circuit that multiplies an input-stream value by each weight of the aperture function and provides products to individual ones of the compositor circuits for further processing.

Ordered traversal of input values plane by plane, row by row, column by column

Ordering input values from a first input point across columns and down rows in a first plane, then continuing across subsequent planes up to a last plane.

Finalization posting output values in order after complete composition per position

The finalization circuit posts output values to the output port in order after receiving and processing a complete composition for each specific position in the P by R by C array of input values.

Simultaneous operation producing output while inputs are received

All circuitry is active simultaneously and the output stream is produced while inputs are being received.

Constraint that C is not an integral multiple of A

A quantitative constraint that C is not an integral multiple of A.

Overall, the claim set covers a streaming IC with ordered input/output streams, an array of compositor circuits decomposing an L×M×N aperture function, dedicated pathways with delay circuits to manage downstream timing, a control circuit that drives counters and control signals, and a finalization circuit that posts output values in order. Dependent refinements include mass multiplication by aperture weights, detailed input ordering across 3D planes, completion-based ordered posting, simultaneous operation with output production during input reception, and a stated constraint relating C and A.

Stated Advantages

Enables computing products common-input-by-all-weights via mass multipliers and decomposed sub-functions.

Supports streaming/ordered processing that avoids redundant memory reads as described in the provided content.

Avoids RAM buffering as described in the provided content.

Enables N-up parallel processing with repackaging/buffering for throughput alignment as described in the provided content.

Documented Applications

Streaming computation of convolutional nodes using an everted aperture-function decomposition, including 3D convolution examples such as 3 by 3 by 3 and 7 by 7 by 7.

Extension to 3D voxel (3 by 3 by 3) convolution with plane buffers and edge-case logic for first/last rows/columns/planes as described in the provided content.

Application of the architecture to MaxPool, Concatenation, Dense node, Global Average, Local Average, and Crop/Subset.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.