Methods for detecting copy-number variations in next-generation sequencing

Inventors

Ivanov, DmitriXU, Zhenyu

Assignees

Sophia Genetics SA

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-12731658-B2

Patent

Publication Date

2026-09-08

Expiration Date


Abstract

Copy Number Variants (CNV) detection methods described herein may efficiently integrate CNV detection into the workflow for a next generation sequencer (NGS) data processing, in parallel with SNP and INDEL variant calling. CNV detection methods as described herein may be performed by analyzing the coverage pattern across a suitable set of genomic regions or amplicons and across a batch of samples from different patients. The proposed methods do not require the use of specifically chosen reference samples as inputs to the workflow, but rather automatically select a set of reference samples from the same batch, for each sample being tested. The CNV detection methods may reliably detect CNVs in a set of samples without prior assumptions about the CNV status of any of those samples. Embodiments described herein may also apply the CNV detection scheme iteratively to further improve the detection performance, especially in the case of more frequent CNV occurrence. Since the knowledge on the CNVs in reference samples may improve their comparison with the sample being tested, the proposed methods may further comprise the step of iteratively feeding back the information about the CNVs found in the samples from any detection step into the next iteration step. The proposed methods may also further use additional information available from the NGS workflow about the samples, such as information on SNP fractions, as input to the NGS CNV detection.

Core Innovation

The invention provides a method for detecting copy-number values (CNV) by integrating CNV detection into a single, targeted high-throughput sequencing experiment. A pool of DNA samples is enriched with a target enrichment technology, where each enriched DNA sample is associated with a library of pooled fragments from a set of amplicons/regions. The enriched amplicons/regions are sequenced with a high-throughput sequencer to generate raw sequencing data for analysis.

The method analyzes raw sequencing data with a genomic data analyzer by cleaning sequencing data to remove low-quality bases and adapter sequences and aligning cleaned reads to a reference genome. It generates, from aligned reads, a coverage count for each sample and for each amplicon/region. It then repeatedly normalizes coverage counts based on a prior estimate of copy-number values for each sample, where the prior estimate is initialized at the first iteration and updated from previous iterations.

For each sample, the method automatically selects a set of reference samples whose normalized coverage patterns are closest to the normalized coverage pattern of the current sample, without requiring specifically chosen, dedicated control samples. The method further normalizes coverage counts for each amplicon/region based on the normalized coverage count of the current sample and the normalized coverage count of the selected reference samples. Copy-number values are then estimated for each sample using a Hidden Markov Model (HMM) as a function of coverage counts in the sample and in the selected reference samples, while utilizing the estimate of copy-number values calculated over previous iterations, and the iterations stop when estimates converge, reach a cycle, or a pre-defined limit is reached.

Claims Coverage

The independent claim is directed to a method that integrates CNV detection into a single targeted high-throughput sequencing experiment and performs iterative CNV estimation using reference-sample selection without dedicated control samples and an HMM over amplicons/regions. The relevant dependent claims refine the HMM scoring/likelihood and confidence output, add a confidence-threshold exclusion rule, specify a quantitative rule for reference-sample set size, and add optional filtering and additional inputs.

Integrated targeted high-throughput sequencing CNV detection with iterative estimation

A method for detecting copy-number values (CNV), wherein detection of CNVs is integrated into a single, targeted high-throughput sequencing experiment, comprising enriching a pool of DNA samples with a target enrichment technology; sequencing each amplicon/region with a high-throughput sequencer; and analyzing raw sequencing data with a genomic data analyzer to determine CNVs for each sample and each amplicon/region, including generating coverage counts and repeating over a plurality of iterations to estimate copy-number values via normalization, automatic reference selection without dedicated control samples, amplicon/region normalization, and HMM-based estimation with a stop condition (convergence, cycle, or pre-defined iteration limit).

Reference sample selection without dedicated controls via closest normalized coverage patterns

Selecting, automatically, for each sample, a set of reference samples as the samples with the closest normalized coverage count to the normalized coverage count of the current sample, where the selecting does not require specifically chosen, dedicated control samples, and the closest coverage pattern is selected by calculating for each sample a distance from a current sample, sorting by increasing distance, and choosing reference samples from the top of the order having the smallest distances.

HMM-based copy-number estimation using sample and reference normalized coverage across iterations

For each sample, estimating, using a Hidden Markov Model (HMM), the copy-number values in the sample as a function of at least the coverage counts in the sample and of at least the coverage counts in the selected set of reference samples for the sample, utilizing the estimate of the copy-number values calculated over previous iterations, with iteration stopping based on convergence, cycle, or reaching a pre-defined limit.

HMM score with likelihoods, non-normal penalties, neighbor-transition penalties and forward-backward confidence

Computing log-likelihoods for each assumed copy-number value from coverage/reference-normalized coverage and noise, defining an HMM score with non-normal and neighbor-transition penalties, and using a forward-backward algorithm to find CNV states that minimize the HMM score while deriving per-amplicon/region confidence as the minimal possible HMM-score increase.

Confidence-threshold exclusion of low-confidence copy-number values

Excluding each of the copy-number values with the confidence level for each amplicon/region below a threshold.

Quantitative rule for reference-sample set size

Determining the number of reference samples NR in each set according to NR=[0.25*N]+2, where N is the total number of samples.

Principal-component filtering of coverage counts

Applying a principal-component filter to the coverage count generated for each sample and for each amplicon/region.

SNP fractions used with coverage to estimate copy-number values

Calculating the estimate of the copy-number values using information on the single-nucleotide polymorphisms (SNP) fractions and coverage, wherein a percentage of SNP fractions is indicative of a duplication.

Across the independent claim, the main inventive elements are iterative normalization and copy-number estimation, automatic reference-sample selection based on closest normalized coverage patterns without dedicated controls, and HMM-based inference of copy-number values per amplicon/region, with specific dependent refinements including HMM score/likelihood formulation with penalties, confidence derivation and threshold exclusion, quantified reference-set sizing, principal-component filtering, and use of SNP fractions with coverage.

Stated Advantages

Not requiring specifically chosen, dedicated control samples for selecting reference samples.

Stopping the iteration and outputting inferred copy-number values when estimates converge, reach a cycle, or a pre-defined limit is reached.

Providing confidence levels per amplicon/region for inferred copy-number states.

Enabling experimental performance comparison to MLPA with detection of CNVs and a low-quality/false-positive rate [only as described in the provided partial content summary].

Documented Applications

Detecting copy-number variants (CNVs) across a set of amplicons/regions in targeted high-throughput sequencing experiments, with experimental comparison to MLPA [only as described in the provided partial content summary].

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.