System and method for management of compressed sequencing files
Inventors
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
Systems and methods for management of storing and analyzing genetic sequencing data. In some embodiments disclosed herein, a method for converting a compressed SAM file back into a raw FASTQ file, wherein the information of the raw FASTQ file is substantively identical to that which was stored in the original FASTQ file from which the compressed SAM file is based is provided. The method advantageously enables storage of the smaller compressed SAM files for reliable, efficient reconstruction of the original FASTQ file when needed.
Core Innovation
A method regenerates a FASTQ file from a compressed sequence alignment map file. The compressed sequence alignment map file includes alignment data for a plurality of sequence strings for more than one aligned reads of clusters on a flow cell. The method is based on determining an fqsum of an original FASTQ file, where each entry of the original FASTQ file is defined by at least four line-separated fields per sequence and the fqsum represents an order invariant checksum of the original FASTQ file.
Base-call quality information is preserved and used to reconstruct original base quality scores. The method stores a quality score of each base call for individual instances of the plurality of sequence strings, and reconstructs original base quality scores by at least one of retrieving the stored quality score or applying an inverse of a model of the stored quality score. To control the alignment data used for regeneration, the compressed sequence alignment map file is sorted, and primary alignments are extracted.
The regeneration writes FASTQ content and ensures correspondence with the original FASTQ. The method ensures all FASTQ records from the original FASTQ file are stored in the compressed SAM file, reconstructs original base quality scores, and writes at least three of the four line-separated fields per sequence of the compressed sequence alignment map file to a regenerated FASTQ file. The regenerated FASTQ file is then compressed, and the order-invariant fqsum supports integrity verification of regenerated output identity.
Claims Coverage
The provided document includes three independent claims: clm-00001, clm-00011, and clm-00020. Each independent claim covers a regeneration workflow that combines an order invariant fqsum determined from an original FASTQ, base-call quality score storage and reconstruction, primary-alignment extraction after sorting, writing selected FASTQ fields, and compressing the regenerated FASTQ. The main differences among the independent claims relate to sorting basis and the specific claim coverage includes a computer program product version of the same workflow.
Order-invariant fqsum determination for original FASTQ
Determining an fqsum of an original FASTQ file on which the compressed sequence alignment file is based, wherein each entry of the original FASTQ file is defined by at least four line-separated fields per sequence and the fqsum represents an order invariant checksum of the original FASTQ file.
Stored per-base quality score and reconstruction via retrieval or inverse modeling
Storing a quality score of each base call for individual instances of the plurality of sequence strings; and reconstructing original base quality scores by at least one of retrieving the stored quality score or applying an inverse of a model of the stored quality score.
Sorting and extracting primary alignments from compressed sequence alignment map
Sorting the compressed sequence alignment map file by one or more reference genome coordinates or by sequence string; and extracting primary alignments from the compressed sequence alignment map file.
Ensuring complete FASTQ record correspondence and regenerating selected FASTQ fields
Ensuring all FASTQ records from the original FASTQ file are stored in the compressed SAM file; writing at least three of the four line-separated fields per sequence of the compressed sequence alignment map file to a regenerated FASTQ file; and compressing the regenerated FASTQ file.
Computer program product for regenerating FASTQ from compressed sequence alignment map
A computer program product encoded on one or more machine-readable storage media and comprising instructions for determining an fqsum, storing a quality score of each base call, sorting the compressed sequence alignment map file, extracting primary alignments, ensuring all FASTQ records are stored, reconstructing original base quality scores by retrieval or inverse modeling, writing at least three of the four FASTQ line-separated fields, and compressing the regenerated FASTQ file.
Across the independent claims, the core claim coverage is regeneration of a FASTQ file from a compressed sequence alignment map using an order invariant fqsum determined from the original FASTQ, storing per-base quality score information, reconstructing original base qualities by retrieval or inverse modeling, sorting the compressed alignment map and extracting primary alignments, ensuring complete FASTQ record correspondence, writing selected FASTQ fields, and compressing the regenerated FASTQ. The computer program product claim provides the same workflow as instructions on machine-readable storage media.
Stated Advantages
Supports integrity verification of regenerated output identity using the order-invariant fqsum.
Enables regeneration of a FASTQ file from a compressed sequence alignment map file.
Ensures original base quality scores are reconstructed by retrieving stored quality scores or applying an inverse of a model of the stored quality score.
Documented Applications
Regenerating a FASTQ file from a compressed sequence alignment map file that includes alignment data for reads of clusters on a flow cell.
Operating as a computer program product that regenerates a FASTQ file from a compressed sequence alignment map file.
Interested in licensing this patent?