Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-7543119-B2

Patent

Publication Date

2009-06-02

Expiration Date


Abstract

A vector processing system provides high performance vector processing using a System-On-a-Chip (SOC) implementation technique. One or more scalar processors (or cores) operate in conjunction with a vector processor, and the processors collectively share access to a plurality of memory interfaces coupled to Dynamic Random Access read/write Memories (DRAMs). In typical embodiments the vector processor operates as a slave to the scalar processors, executing computationally intensive Single Instruction Multiple Data (SIMD) codes in response to commands received from the scalar processors. The vector processor implements a vector processing Instruction Set Architecture (ISA) including machine state, instruction set, exception model, and memory model.

Core Innovation

The invention describes a Vector Processing System (VPS) SoC implementation in which scalar cores drive a slave vector processor that executes SIMD/vector ISA instruction blocks. The vector processor includes Functional Lane execution pipelines with instruction control/status, instruction translation and caching, vector register structures, and a defined machine state, exception model, and memory model.

The memory model is coherent with the scalar cache but not coherent among vector operations, enabling defined behavior across scalar and vector domains. The invention provides a Memory Buffer Switch and Memory Lane subsystem to route memory requests from the vector execution pipelines, including request routing and switching across Memory Lanes with blocking and scheduling, coalescing of load/store requests into larger DRAM bursts, and buffering of read returns and instruction-fetch returns.

Credit-based flow control and synchronization instructions are included to enforce ordering where required, including handling of synchronization effects across vector memory operations. The disclosed design also includes SMP support via coherent HyperTransport and northbridge functions, including coherency probe filtering to reduce traffic, together with instruction-fetch return handling and vector register/control structures to coordinate execution state, masking, predication, and control/status behavior.

Claims Coverage

The provided document content includes two independent claims. Across both independent claims, the inventive features focus on concurrent floating point execution of vector-instruction parts, out-of-order completion across execution units, consolidation of multiple memory requests into a single memory access via a memory buffer switch, and processing of those requests according to an x86-implemented coherency domain.

Concurrent parts execution on multiple floating point execution units for a vector instruction stream

An instruction control unit configured to control the floating point execution units according to a stream of vector instructions, wherein the multiple floating point execution units operate concurrently on parts of the same vector instruction; and wherein the parts of multiple vector instructions are executed independently by the separate floating point execution units.

Out-of-order completion across floating point execution units

The parts of the vector instructions are completed out of order with respect to the parts in other floating point execution units.

x86-compatible processor interface receiving the stream of vector instructions

A processor interface compatible with an x86 processor and configured to receive the stream of vector instructions from the x86 processor, wherein the x86 processor controls the plurality of floating point units through a compatible interface.

Shared access to coherent memory and x86 coherency-domain memory processing

The x86 processor, and the plurality of floating point units share access to coherent memory; and the memory buffer switch is configured to consolidate at least two memory requests from the floating point execution units into a single memory access operation directed to one of the memory channels and the at least two memory requests are processed according to a coherency domain implemented by the x86 processor.

Both independent claims cover a system that executes vector instruction parts on multiple floating point execution units with out-of-order completion across units and uses a memory buffer switch to consolidate multiple memory requests into a single memory access directed to a memory channel. The consolidated requests are processed according to a coherency domain implemented by the x86 processor, and one independent claim additionally requires an x86-compatible processor interface controlling the floating point units and sharing coherent memory access.

Stated Advantages

Documented Applications

No documented applications found

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.