Methods for processing next-generation sequencing genomic data

Inventors

Song, LinSTEIJGER, TamaraBEHR, JonasNOVAK, AdamHernandez, DavidXU, Zhenyu

Assignees

Sophia Genetics SA

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-11923049-B2

Patent

Publication Date

2024-03-05

Expiration Date


Abstract

A genomic data analyzer system method to analyze next generation sequencing genomic data from a sourcing laboratory. The method includes receiving, with a processor, a next generation sequencing analysis request from a sourcing laboratory, the next generation sequencing request comprising at least a raw next generation sequencing data file and the sourcing laboratory identification; identifying, with a processor, a first set of characteristics associated with the next generation sequencing analysis request, the first set of characteristics comprising at least a target enrichment technology identifier, a sequencing technology identifier, and a genomic context identifier; configuring, with a processor, a data alignment module to align the input raw sequencing data file in accordance with at least one characteristic of said first set of characteristics; and aligning, with the data alignment module processor, the input sequencing data to a genomic sequence.

Core Innovation

A computer-implemented method analyzes next generation sequencing genomic data from a sourcing laboratory by receiving, with a processor, via a communication network, a next generation sequencing analysis request generated via a laboratory computing infrastructure of the sourcing laboratory. The next generation sequencing request comprises at least a raw sequencing data file and a sourcing laboratory identification. The method identifies first set of characteristics associated with the next generation sequencing analysis request, including at least a target enrichment technology identifier, a sequencing technology identifier, and a genomic context identifier.

The method configures a data alignment module to align the raw sequencing data file in accordance with the first set of characteristics, aligns the raw sequencing data to a genomic sequence to produce a raw alignment file, and identifies a second set of characteristics associated with the alignment data. The second set of characteristics comprises at least a data alignment pattern identifier. The method configures the data alignment module to refine at least one subset of the input raw sequencing data in accordance with at least one characteristic of the first set of characteristics and at least one characteristic of the second set of characteristics to produce a refined alignment data file.

The refinement comprises re-aligning a subset of the raw alignment data file, and the method configures a variant calling module to identify variants associated with the raw sequencing data file in accordance with at least one characteristic of the first set of characteristics and at least one characteristic of the second set of characteristics. The method identifies genomic variants, with the variant calling module processor, in the refined alignment data file.

Claims Coverage

The independent claims are clm-00001 and clm-00017. Across these claims, the inventive features are centered on deriving request-specific characteristics and using them to configure a data alignment module and a variant calling module, with claim clm-00001 further refining alignments using alignment-pattern characteristics before variant calling.

Receiving sequencing requests with lab identification

Receiving, with a processor, via a communication network, a next generation sequencing analysis request from a sourcing laboratory, the next generation sequencing request comprising at least a raw sequencing data file and a sourcing laboratory identification.

Deriving request characteristics for alignment and refinement

Identifying, with a processor, a first set of characteristics associated with the next generation sequencing analysis request, the first set of characteristics comprising at least a target enrichment technology identifier, a sequencing technology identifier, and a genomic context identifier.

Configuring alignment using request characteristics

Configuring, with a processor, a data alignment module to align the raw sequencing data file in accordance with at least one characteristic of said first set of characteristics, and aligning, with the data alignment module processor, the raw sequencing data to a genomic sequence to produce a raw alignment file.

Extracting alignment-pattern characteristics and refining subsets by re-aligning

Identifying, with a processor, a second set of characteristics associated with the alignment data from the raw alignment data file, the second set of characteristics comprising at least a data alignment pattern identifier; configuring, with a processor, the data alignment module to refine at least one subset of the input raw sequencing data in accordance with at least one characteristic of the first set of characteristics and at least one characteristic of the second set of characteristics; refining, the subset of the raw alignment data to produce a refined alignment data file, wherein refining comprises re-aligning a subset of the raw alignment data file.

Configuring variant calling using request and alignment-pattern characteristics

Configuring, with a processor, a variant calling module to identify variants associated with the raw sequencing data file in accordance with at least one characteristic of said first set of characteristics and at least one characteristic of the second set of characteristics; and identifying genomic variants in the refined alignment data file.

Handling multiple sourcing laboratories with shared characteristic-driven configuration

Receiving, with a processor, via a communication network, a next generation sequencing analysis request from each of the plurality of sourcing laboratories, the next generation sequencing request comprising at least a raw next generation sequencing data file and the sourcing laboratory identification; identifying a first set of characteristics comprising at least a target enrichment technology identifier, a sequencing technology identifier, and a genomic context identifier; configuring a data alignment module to align the input raw sequencing data file in accordance with at least one characteristic of the first set of characteristics; configuring a variant calling module to identify variants associated with the input sequencing data in accordance with at least one characteristic of the first set of characteristics; and identifying genomic variants in the aligned genomic sequence.

Using a laboratory information database for laboratory characteristics

Appending, for each sourcing laboratory, the target enrichment technology identifier, the sequencing technology identifier, and the genomic context identifier to a laboratory information database via the processor.

Updating laboratory information database on registration and changes

Updating the laboratory information database, upon registration of a new sourcing laboratory and upon update of an already registered sourcing laboratory.

Overall, the claims cover configuring alignment and variant calling modules based on request-specific identifiers (target enrichment technology identifier, sequencing technology identifier, genomic context identifier) and, in clm-00001, alignment-derived data alignment pattern identifier, including refining alignments by re-aligning subsets before identifying genomic variants.

Stated Advantages

Improved specificity by removing false positives via optimized handling of challenging alignment contexts.

Scalable and backward/forward compatible onboarding of new laboratories without manual per-lab configuration.

Documented Applications

Multi-laboratory workflows for analyzing next generation sequencing genomic data received from sourcing laboratories.

Analyzing genomic data produced using Illumina MiSeq and Ion Torrent PGM, and using amplicon-based and probe-based enrichment (e.g., indicated DNA enrichment assay/kits).

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.