Bioinformatics systems, apparatuses, and methods executed on an integrated circuit processing platform
Inventors
Van Rooyen, Pieter • Ruehle, Michael • Mehio, Rami
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
A system, method and apparatus for executing a bioinformatics analysis on genetic sequence data includes an integrated circuit formed of a set of hardwired digital logic circuits that are interconnected by physical electrical interconnects. One of the physical electrical interconnects forms an input to the integrated circuit that may be connected with an electronic data source for receiving reads of genomic data. The hardwired digital logic circuits may be arranged as a set of processing engines, each processing engine being formed of a subset of the hardwired digital logic circuits to perform one or more steps in the bioinformatics analysis on the reads of genomic data. Each subset of the hardwired digital logic circuits may be formed in a wired configuration to perform the one or more steps in the bioinformatics analysis.
Core Innovation
The invention provides a genomics analysis platform for executing a sequence analysis pipeline using one or more first integrated circuits and one or more second integrated circuits. The first integrated circuits form processing units responsive to one or more software algorithms, and the second integrated circuits form programmable logic devices with hardware logic circuits arranged as processing engines to perform one or more different genomic processing steps of the sequence analysis pipeline.
A shared memory stores genetic sequence data, and both the first integrated circuits and the second integrated circuits access the genetic sequence data in the shared memory based on one or more cache coherency protocols. The platform uses heterogeneous integrated-circuit execution of genomic processing steps with shared-memory access to genetic sequence data.
In the described FPGA/ASIC-based approach, programmable logic performs haplotype-based pair-HMM computation for variant calling using match/insert/delete dynamic programming states. The computation is organized using swath or wavefront diagonal traversal with pipelined latencies and limited state storage per swath boundary, while transitioning probability and prior generation is driven by lookup tables.
System-level emphasis is placed on throughput and reducing data movement by using shared-memory architectures and cache coherency between CPU/FPGA components rather than relying on PCIe-based loose integration for transferring data. The described execution includes mechanisms to support out-of-order hardware execution by associating job results to records, and optionally uses log-domain arithmetic to handle dynamic range in the HMM computations.
Claims Coverage
The independent claim is directed to a genomics analysis platform architecture that combines first integrated circuits and programmable-logic devices accessing genetic sequence data via shared memory using cache coherency protocols. The supported dependent claims refine how cache coherency protocols govern alternating processing and add example modeling and implementation details.
Heterogeneous integrated-circuit execution of a sequence analysis pipeline
One or more first integrated circuits form processing units responsive to one or more software algorithms to perform one or more genomic processing steps of a sequence analysis pipeline, and one or more second integrated circuits form programmable logic devices whose hardwired digital logic circuits are arranged as processing engines to perform one or more different genomic processing steps of the sequence analysis pipeline.
Shared-memory genetic sequence access governed by cache coherency protocols
A shared memory is electronically connected to the one or more first integrated circuits via at least a portion of physical electronic interconnects, stores genetic sequence data, and is accessed by each of the one or more first integrated circuits and each of the one or more second integrated circuits based on one or more cache coherency protocols.
Cache coherency governed alternation based on computational-intensity level
Cache coherency protocols govern alternating processing of stored genetic sequence data in shared memory according to a computational-intensity level associated with a portion of that data.
Pair hidden Markov model estimating probability for an accelerated subset
A subset of discrete operations that meet a computational-intensity threshold uses a pair hidden Markov model to estimate the probability of observing a particular read based on a candidate haplotype.
Programmable logic circuit implemented as FPGA or ASIC
A programmable logic circuit is implemented as a field-programmable gate array or an Application Specific Integrated Circuit.
The claims coverage centers on executing a sequence analysis pipeline using both software-instructed processing units and programmable-logic processing engines, while accessing genetic sequence data in shared memory through cache coherency protocols. Dependent refinements describe alternating processing based on computational-intensity levels and include pair hidden Markov model probability estimation and FPGA or ASIC implementation details.
Stated Advantages
Enables throughput-oriented execution by using swath or wavefront processing with pipelined latencies and limited state storage per swath boundary.
Improves system efficiency by reducing data movement through shared-memory cache-coherent CPU/FPGA cooperation rather than PCIe-based loose integration.
Supports handling of HMM computation dynamic range via optional log-domain arithmetic.
Documented Applications
Haplotype-based variant calling using pair-HMM computation organized with match/insert/delete states in programmable logic.
Genomic pipeline execution in which mapping/alignment and variant calling are performed via CPU/FPGA cooperation within shared-memory architectures.
Interested in licensing this patent?