Methods for detecting variants in next-generation sequencing genomic data
Inventors
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
A genomic data analyzer may be configured to detect and characterize, with a variant calling module, genomic variant scenarios in sequencing reads from an enriched patient genomic sample comprising a combination of a first repeat pattern and a second repeat pattern, such as repeats of homopolymer (single nucleotide) and/or heteropolymer (multiple nucleotide) basic motifs. The variant calling module may estimate the probability distribution of the length of the first repeat pattern and the probability distribution of the length of the second repeat pattern by comparing the distribution of the repeat pattern length measurements in patient data to the distribution of the repeat pattern length measurements in control data, in order to remove biases possibly induced by the next generation sequencing laboratory setup both in control and patient data. The variant calling module may further measure, read by read, the joint probability distribution for the first and the second repeat patterns lengths, and compare it with the expected joint probability distribution for various genomic variant scenarios for the patient, each variant scenario being characterized by a first length of the first repeat pattern and a second length of the second repeat pattern, to select the most likely patient genomic variant scenario as the scenario for which the measured joint probability distribution best matches the expected joint probability distribution.
Core Innovation
The invention provides a method for detecting and characterizing a genomic variant scenario as a combination of at least two genomic sequence variants in poly-T and poly-TG tracts of CFTR gene alleles from a cystic fibrosis patient sample. The method uses next generation sequencing data obtained from an enriched genomic sample of the patient and an enriched genomic control sample, and targets and enriches sub-regions corresponding to the poly-T and poly-TG tracts of CFTR gene alleles to generate patient data sequence reads and control data sequence reads.
A central feature is the measurement and comparison of discrete probability distributions of repeat pattern lengths. The method measures probability distributions of the length of poly-TG and poly-T repeat patterns in the control data sequence reads, estimates expected probability distributions for candidate genomic variants relative to the control data, measures patient probability distributions, and selects the estimated patient probability distribution that results in the closest comparison.
The method further estimates expected joint probability distributions for combinations of poly-TG and poly-T allele repeat lengths under genomic variant scenarios defined by alleles being on the same reads. It measures, read by read, the patient joint probability distribution for the length of the poly-TG repeat pattern and the length of the poly-T repeat pattern, compares it with the expected joint probability distributions, and selects the genomic variant scenario that results in the closest comparison.
Claims Coverage
The provided excerpt identifies one independent claim. The independent claim includes at least eleven core inventive processing features: enriched patient and control sampling, repeat-length probability distribution measurement, variant-dependent expected distribution estimation, closest-match selection for each repeat tract, and scenario selection using measured read-by-read joint probability distributions for allele-linked scenarios.
Enriched patient targeting of poly-T and poly-TG tracts
Obtain patient data sequence reads from an enriched genomic sample of the patient using next generation sequencing by targeting and enriching sub-regions corresponding to the poly-T and poly-TG tracts of CFTR gene alleles.
Enriched control targeting of poly-T and poly-TG tracts
Obtain control data sequence reads from an enriched genomic control sample using next generation sequencing by targeting and enriching sub-regions corresponding to the poly-T and poly-TG tracts of CFTR gene alleles.
Control poly-TG repeat probability distribution measurement
Measure a probability distribution of the length of a poly-TG repeat pattern in the plurality of control data sequence reads.
Expected poly-TG distribution estimation for candidate variants
For each possible genomic variant of the poly-TG repeat pattern relative to the control data poly-TG repeat pattern, estimate the expected probability distribution of the length of the poly-TG repeat pattern as a function of the genomic variant and of the measured probability distribution in the control data sequence reads.
Closest-match selection for poly-TG variant repeat distribution
Measure a patient probability distribution of the length of the poly-TG repeat pattern and compare it with the expected probability distributions; for each possible genomic variant, compare the measured patient probability distribution and the expected probability distribution, then select the estimated patient probability distribution that results in the closest comparison.
Control poly-T repeat probability distribution measurement
Measure a probability distribution of the length of a poly-T repeat pattern in the plurality of control data sequence reads as the expected probability distribution length of the poly-T repeat pattern in the control sample.
Expected poly-T distribution estimation for candidate variants
For each possible genomic variant of the poly-T repeat pattern relative to the control data poly-T repeat pattern, estimate the expected probability distribution of the length of the poly-T repeat pattern as a function of the genomic variant and of the measured probability distribution of the length of the poly-T repeat pattern in the plurality of control data sequence reads.
Closest-match selection for poly-T variant repeat distribution
Measure a patient probability distribution of the length of the poly-T repeat pattern, compare it with the expected probability distributions for each possible genomic variant, and select the estimated patient probability distribution that results in the closest comparison.
Scenario-defined expected joint distribution estimation (alleles on same reads)
Estimate a first expected joint probability distribution of the length of the poly-TG repeat pattern and of the length of the poly-T repeat pattern for at least one first genomic variant scenario where a first allele selected in the poly-TG selection and a first allele selected in the poly-T selection are on the same reads, while a second allele selected in the poly-TG selection and a second allele selected in the poly-T selection are on the same reads; estimate a second expected joint probability distribution for at least one second genomic variant scenario where the allele pairing on the same reads is swapped between the first and second alleles.
Read-by-read measurement of patient joint distribution
Measure, read by read, the patient joint probability distribution for the length of the poly-TG repeat pattern and the length of the poly-T repeat pattern in the plurality of patient data sequence reads.
Closest-match scenario selection from joint distribution comparisons
Compare the measured patient joint probability distribution with the first expected joint probability distribution and with the second expected joint probability distribution for the corresponding genomic variant scenarios, and select the genomic variant scenario characterizing the actual genomic variant scenario for the patient as the scenario which results in the closest comparison.
The independent claim is directed to selecting a cystic fibrosis CFTR poly-T/poly-TG genomic variant scenario by estimating expected repeat-length probability distributions from control data, selecting closest-match expected distributions for each repeat tract, and then resolving allele-linked scenario combinations using measured read-by-read joint probability distributions and closest-comparison selection.
Stated Advantages
Improved analytical sensitivity and analytical specificity versus an FDA-approved reference and validator comparisons, as reported in Italian cohorts.
Repeatability and reproducibility.
Documented Applications
Diagnostic or neonatal and carrier screening, assisted reproduction, and genetic counseling.
Implementation and testing on CFTR assays including Multiplicom CFTR MASTR Dx, drMID Dx, Devyser CFTR kit, and Sophia Genetics DDM.
Use of performance reporting based on Italian cohorts and cohorts using a Sophia Genetics DDM implementation on the specified commercial CFTR assays.
Adaptation to alternative CFTR kits and potential broader applicability to other repeat-associated diseases.
Interested in licensing this patent?