Copy number variant caller

Inventors

HAAS, Kevin R.Wang, XinGRAUMAN, Peter V.

Assignees

Myriad Genetics IncMyriad Womens Health Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-11232850-B2

Patent

Publication Date

2022-01-25

Expiration Date


Abstract

Direct targeted sequencing (DTS) methods and a hidden Markov model (HMM) can be used to call the copy number of a segment of interest within a region of interest. Described herein are methods for calling a copy number variant or a copy number variant abnormality using an HMM, and methods for determining a copy number based on a copy number likelihood model, in a test sequencing library that has be sequenced using DTS methods. Also described herein are methods for determining a copy number of a segment, including accounting for spurious capture probes that may arise from the DTS methods.

Core Innovation

The invention relates to a copy number variant caller that determines a copy number of an interrogated segment of nucleic acids within a region of interest of a genome. Sequencing reads from a test sequencing library are mapped to the interrogated segment, and the number of sequencing reads mapped to the interrogated segment is determined. A copy number likelihood model is determined using a plurality of likelihood distributions associated with an expected number of sequencing reads mapped to the interrogated segment.

A hidden Markov model is built with one or more hidden states comprising a copy number corresponding to the interrogated segment or a plurality of sub-segments within the interrogated segment. An observation state is included that comprises the number of sequencing reads mapped to the interrogated segment, and the observation state is linked to the copy number likelihood model. The hidden Markov model is parameterized by adjusting the copy number likelihood model to fit the determined number of sequencing reads mapped to the interrogated segment by allowing portions of the likelihood distributions to float.

The parameterized hidden Markov model is then used to determine the most probable copy number of the interrogated segment. In additional embodiments described, likelihood distributions include negative binomial distributions that are not Poisson distributions, and the modeling accounts for GC content bias and spurious capture probes using a spurious capture probe indicator. The approach further accounts for noise in mapped read counts by adjusting dispersion and by threshold-based non-calling when noise is above a predetermined threshold, and it can improve detection of CNVs spanning multiple segments.

Claims Coverage

The independent claim provides a method-level coverage with five inventive features centered on constructing a copy number likelihood model from expected mapped read counts and embedding that model into a hidden Markov model that is parameterized and optimized to output the most probable copy number for the interrogated segment.

Mapping reads to the interrogated segment and counting mapped reads

Mapping a plurality of sequencing reads of nucleic acids generated from a test sequencing library to the interrogated segment; determining a number of sequencing reads mapped to the interrogated segment.

Constructing a copy number likelihood model from expected mapped read counts

Determining a copy number likelihood model using a plurality of likelihood distributions associated with an expected number of sequencing reads mapped to the interrogated segment.

Building a hidden Markov model with hidden copy-number states and a read-count observation state

Building a hidden Markov model comprising one or more hidden states comprising a copy number corresponding to the interrogated segment or a plurality of sub-segments within the interrogated segment; an observation state comprising the number of sequencing reads mapped to the interrogated segment; and the copy number likelihood model.

Parameterizing the hidden Markov model by fitting the likelihood distributions to observed mapped read counts

Parameterizing the hidden Markov model by adjusting the copy number likelihood model to fit the determined number of sequencing reads mapped to the interrogated segment by allowing portions of the likelihood distributions to float.

Determining the most probable copy number by optimizing the parameterized HMM

Determining a most probable copy number of the interrogated segment by optimizing the parameterized hidden Markov model.

Across the independent claim, the coverage is directed to an HMM-based copy number determination method where a copy number likelihood model drives an HMM with hidden copy-number states and a read-count observation state, followed by parameterization by fitting with floating likelihood portions and then optimization to output the most probable copy number.

Stated Advantages

Reduced false positives and better detection of true CNVs spanning multiple segments.

Documented Applications

Copy number variant calling and determining copy number of an interrogated segment within a region of interest of a genome.

Embodiments associated with cell-free DNA, including fetal ctDNA and circulating tumor cfDNA.

Medical diagnosis and reporting.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.