Methods and systems for genetic analysis

Inventors

Bartha, Gabor T. • Chandratillake, Gemma • Chen, Richard • Garcia, Sarah • Lam, Hugo Yu Kor • Luo, Shujun • Pratt, Mark R. • West, John

Assignees

Personalis Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-11649499-B2

Patent

Publication Date

2023-05-16

Expiration Date


Abstract

This disclosure provides systems and methods for sample processing and data analysis. Sample processing may include nucleic acid sample processing and subsequent sequencing. Some or all of a nucleic acid sample may be sequenced to provide sequence information, which may be stored or otherwise maintained in an electronic storage location. The sequence information may be analyzed with the aid of a computer processor, and the analyzed sequence information may be stored in an electronic storage location that may include a pool or collection of sequence information and analyzed sequence information generated from the nucleic acid sample. Methods and systems of the present disclosure can be used, for example, for the analysis of a nucleic acid sample, for producing one or more libraries, and for producing biomedical reports. Methods and systems of the disclosure can aid in the diagnosis, monitoring, treatment, and prevention of one or more diseases and conditions.

Core Innovation

The invention is a method for analyzing nucleic acid sample(s) of a subject by generating at least a first subset and a second subset of nucleic acid molecules from one or more nucleic acid samples. The first subset is selectively enriched with probes that selectively enrich for a genomic feature consisting of phased variants, while the second subset is selectively enriched with probes that selectively enrich for a genomic feature consisting of polymorphisms.

The method subjects the selectively enriched first subset to a first assay to yield a first result comprising a first nucleic acid sequence, and subjects the selectively enriched second subset to a second assay to yield a second result comprising a second nucleic acid sequence. The two results are then combined with the aid of a computer processor to generate an output comprising a consensus sequence derived from the first nucleic acid sequence and the second nucleic acid sequence.

The described approaches are supported by workflows that generate genomic-content subsets for targeted enrichment, including Exome Supplement Plus (ESP), high-GC content pulldown (HGCP), and long/region-specific pulldown (LRP). The document further describes recovering sequence content missed by standard exome workflows through multipronged subset generation and intersection/product comparisons, including ACE/Exome+.

Claims Coverage

The document provides one independent method claim focused on dual-subset selective enrichment for phased variants and polymorphisms, separate assays to generate sequences, and computer-assisted combining into a consensus sequence. Additional inventive limitations include combining strategies based on precedence rules and/or quality and read coverage metrics to resolve discordances.

Dual selective enrichment for phased variants and polymorphisms

Generating at least a first subset and a second subset of nucleic acid molecules from one or more nucleic acid samples, wherein the first subset is selectively enriched with probes selectively enriching genomic features consisting of phased variants, and the second subset is selectively enriched with probes selectively enriching genomic features consisting of polymorphisms.

Assay-derived sequences from enriched subsets

Subjecting the first subset to a first assay to yield a first result comprising a first nucleic acid sequence, and subjecting the second subset to a second assay to yield a second result comprising a second nucleic acid sequence.

Computer-assisted consensus combining of sequences

Combining, with the aid of a computer processor, the first result and the second result to generate an output comprising a consensus sequence from the first nucleic acid sequence and the second nucleic acid sequence.

Consensus resolution using precedence rules or genomic context

Combining the first and second nucleic acid sequences using a precedence rule based on genomic context(s) and/or assay(s) to resolve disagreements between multiple sequencing data sets.

Consensus resolution using quality and read coverage metrics

Combining two nucleic acid sequences using quality metrics and read coverage metrics to resolve discordant genotypes.

Overall, claim coverage centers on generating two selectively enriched nucleic acid subsets for different genomic feature categories, producing sequence results by separate assays, and using a computer processor to combine them into a consensus sequence, with discordance resolution using precedence rules and/or quality and read coverage metrics.

Stated Advantages

Recovering sequence content missed by standard exome workflows.

Documented Applications

Disease diagnosis, monitoring, treatment, and/or prevention using analyzed outputs and biomedical reports generated from genetic analysis results.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.