FastQ/FastA compression systems and methods
Inventors
Nazari, Foad • Patel, Sneh • Murray, Emma K. • Schena, Giana J.
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
Systems and methods to analyze and significantly compress FastQ and/or FastA datasets are disclosed. The methodology includes algorithms to compress sequences, quality scores and identifiers of read files. The method relies on reducing the dimension and redundancy in genomic data in a unique and optimal way and in the binary format. The methodology also includes the decoding protocols to decompress the compressed data with zero loss.
Core Innovation
The disclosed invention relates to data compression of genomic data by receiving a data file comprised of sequence bases, quality scores, and identifiers, where the sequence bases comprise regular bases adenine (A), cytosine (C), guanine (G), and thymine (T) and irregular bases. The method applies an optimization algorithm to the quality scores, sequence k-mers from the sequence bases, and the identifiers, and splits long reads of the sequence bases and the quality scores into smaller segments.
After read splitting and deletion of duplicated and semi-duplicated reads, the invention performs dimensionality reduction on the sequence bases. It stores a template of the identifier that is consistent across the data file, and it detects and stores the location and type of each irregular base.
The encoding step represents each regular base by one of four two-digit binary numbers and represents all irregular bases by one of four two-digit binary numbers, while maintaining the irregular-base location and type information. The encoded data file is then compressed, where compression is lossless and reference-free in at least some embodiments described in the document, and the described compression supports FastQ and FastA genomic data file types.
Claims Coverage
The partial content provides two independent claims: one method claim and one non-transitory computer storage medium claim. Each independent claim specifies a chain of inventive features totaling receiving and preprocessing genomic sequence, quality score, and identifier data, dimensionality reduction and identifier templating, irregular-base location and type tracking with binary encoding, and lossless compression.
Receiving genomic data with regular and irregular bases
receiving a data file comprised of sequence bases, quality scores, and identifiers, the sequence bases comprising regular bases and irregular bases, the regular bases comprising adenine (A), cytosine (C), guanine (G), and thymine (T)
Optimizing quality scores, k-mers, and identifiers
applying an optimization algorithm to the quality scores, sequence k-mers from the sequence bases, and the identifiers
Splitting long reads
splitting long reads of the sequence bases and the quality scores into smaller segments
Deleting duplicated and semi-duplicated reads
deleting duplicated and semi-duplicated reads for the sequence bases and the quality scores
Dimensionality reduction on sequence bases
performing a dimensionality reduction on the sequence bases
Consistent identifier template
storing a template of the identifier that is consistent across the data file
Irregular base location and type detection and storage
detecting and storing the location and type of each irregular base
Binary encoding of regular and irregular bases as four two-digit values each
encoding the data file in a binary format such that each regular base is represented by one of the four two-digit binary numbers and all irregular bases are represented by one of the four two-digit binary numbers
Lossless compression of the encoded data file
compressing the encoded data file
Lossless compression in the medium-executed process
compressing the data file, wherein the compression is lossless
Across both independent claims, the inventive coverage centers on jointly optimizing quality scores, sequence k-mers, and identifiers; preprocessing by splitting long reads and deleting duplicated and semi-duplicated reads; applying dimensionality reduction; using a consistent identifier template; detecting and storing irregular-base location and type; encoding regular and irregular bases into binary representations using four two-digit binary numbers each; and performing lossless compression.
Stated Advantages
Higher compression rate.
Zero-loss decoding.
Dataset-specific tunable protocol.
Support for FastA.
Faster compression.
Documented Applications
Lossless, reference-free compression of FastQ genomic files and FastA genomic files for genomic sequence run data/cluster data.
Interested in licensing this patent?