FastQ/FastA compression systems and methods

Inventors

Nazari, FoadPatel, SnehMurray, Emma K.Schena, Giana J.

Assignees

Rajant Health Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-12287764-B2

Patent

Publication Date

2025-04-29

Expiration Date


Abstract

Systems and methods to analyze and significantly compress FastQ and/or FastA datasets are disclosed. The methodology includes algorithms to compress sequences, quality scores and identifiers of read files. The method relies on reducing the dimension and redundancy in genomic data in a unique and optimal way and in the binary format. The methodology also includes the decoding protocols to decompress the compressed data with zero loss.

Core Innovation

The disclosed invention relates to data compression of genomic data by receiving a data file comprised of sequence bases, quality scores, and identifiers, where the sequence bases comprise regular bases adenine (A), cytosine (C), guanine (G), and thymine (T) and irregular bases. The method applies an optimization algorithm to the quality scores, sequence k-mers from the sequence bases, and the identifiers, and splits long reads of the sequence bases and the quality scores into smaller segments.

After read splitting and deletion of duplicated and semi-duplicated reads, the invention performs dimensionality reduction on the sequence bases. It stores a template of the identifier that is consistent across the data file, and it detects and stores the location and type of each irregular base.

The encoding step represents each regular base by one of four two-digit binary numbers and represents all irregular bases by one of four two-digit binary numbers, while maintaining the irregular-base location and type information. The encoded data file is then compressed, where compression is lossless and reference-free in at least some embodiments described in the document, and the described compression supports FastQ and FastA genomic data file types.

Claims Coverage

The partial content provides two independent claims: one method claim and one non-transitory computer storage medium claim. Each independent claim specifies a chain of inventive features totaling receiving and preprocessing genomic sequence, quality score, and identifier data, dimensionality reduction and identifier templating, irregular-base location and type tracking with binary encoding, and lossless compression.

Receiving genomic data with regular and irregular bases

receiving a data file comprised of sequence bases, quality scores, and identifiers, the sequence bases comprising regular bases and irregular bases, the regular bases comprising adenine (A), cytosine (C), guanine (G), and thymine (T)

Optimizing quality scores, k-mers, and identifiers

applying an optimization algorithm to the quality scores, sequence k-mers from the sequence bases, and the identifiers

Splitting long reads

splitting long reads of the sequence bases and the quality scores into smaller segments

Deleting duplicated and semi-duplicated reads

deleting duplicated and semi-duplicated reads for the sequence bases and the quality scores

Dimensionality reduction on sequence bases

performing a dimensionality reduction on the sequence bases

Consistent identifier template

storing a template of the identifier that is consistent across the data file

Irregular base location and type detection and storage

detecting and storing the location and type of each irregular base

Binary encoding of regular and irregular bases as four two-digit values each

encoding the data file in a binary format such that each regular base is represented by one of the four two-digit binary numbers and all irregular bases are represented by one of the four two-digit binary numbers

Lossless compression of the encoded data file

compressing the encoded data file

Lossless compression in the medium-executed process

compressing the data file, wherein the compression is lossless

Across both independent claims, the inventive coverage centers on jointly optimizing quality scores, sequence k-mers, and identifiers; preprocessing by splitting long reads and deleting duplicated and semi-duplicated reads; applying dimensionality reduction; using a consistent identifier template; detecting and storing irregular-base location and type; encoding regular and irregular bases into binary representations using four two-digit binary numbers each; and performing lossless compression.

Stated Advantages

Higher compression rate.

Zero-loss decoding.

Dataset-specific tunable protocol.

Support for FastA.

Faster compression.

Documented Applications

Lossless, reference-free compression of FastQ genomic files and FastA genomic files for genomic sequence run data/cluster data.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.