Systems and methods for determining effects of therapies and genetic variation on polyadenylation site selection
Inventors
Frey, Brendan • LEUNG, Michael Ka Kit
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
The present disclosure provides systems and methods for determining effects of genetic variants on selection of polyadenylation sites (PAS) during polyadenylation processes. In an aspect, the present disclosure provides a polyadenylation code, a computational model that can predict alternative polyadenylation patterns from transcript sequences. A score can be calculated that describes or corresponds to the strength of a PAS, or the efficiency in which it is recognized by the 3′-end processing machinery. The polyadenylation model may be used, for example, to assess the effects of anti-sense oligonucleotides to alter transcript abundance. As another example, the polyadenylation model may be used to scan the 3′-UTR of a human genome to find potential PAS.
Core Innovation
The disclosed invention provides a machine-learning polyadenylation code that predicts polyadenylation site (PAS) strength and preferences from genomic sequence context, and uses the predicted PAS preferences to determine how genetic variants and antisense oligonucleotides shift PAS usage. The approach is framed within alternative polyadenylation and the operation of 3′-end processing machinery, so that changes in sequence are evaluated in terms of predicted PAS preference outputs.
A preference computation is performed for each candidate polyadenylation site by applying a trained algorithm to polyadenylation feature vectors extracted from nucleotides in genomic sequence regions. The model converts intermediate representations into site preferences p1…pn, using normalization/softmax-like functions, and uses gradient-based learning during training. Variants can be incorporated by computer processing a reference sequence into a variant sequence based on an antisense oligonucleotide, and the resulting preference sets across reference and variant genomic sequences are compared to estimate an effect.
The disclosed framework further supports genome scanning to locate candidate polyadenylation sites and to identify regions for polyadenylation-targeted therapeutics, including in-silico therapy simulation using antisense blocking. Gradient/saliency analyses are used to identify sequence motifs and tissue-specific polyadenylation features, and the patent describes evaluation involving PAS selection, variant pathogenicity near PAS, and genome-wide PAS discovery.
Claims Coverage
The relevant portion provides one independent claim. It includes five core inventive aspects: (1) generating a reference/variant genomic sequence pair based on an antisense oligonucleotide, (2) identifying candidate polyadenylation sites in each genomic sequence, (3) extracting polyadenylation feature vectors for each candidate site from nucleotides in the genomic sequence, (4) applying a trained algorithm to determine site preferences p1…pn, and (5) computer processing preference sets across the reference and variant sequences to determine an effect, followed by administering a therapeutically effective amount that modulates polyadenylation of at least one candidate site.
Antisense-based reference and variant genomic sequences
Providing a plurality of genomic sequences comprising a reference sequence and a variant sequence obtained by computer processing the reference sequence based at least in part on the antisense oligonucleotide, wherein the antisense oligonucleotide is complementary to at least a portion of the reference sequence.
Candidate polyadenylation site identification and feature vector extraction
For each of the plurality of genomic sequences: identifying a plurality of candidate polyadenylation sites in the genomic sequence and extracting a polyadenylation feature vector for each candidate polyadenylation site, wherein each feature vector comprises a set of features determined based at least in part on a set of nucleotides in the genomic sequence.
Trained algorithm for PAS preference determination
Applying a trained algorithm to the plurality of polyadenylation feature vectors to determine a set of preferences p1, p2, . . . , pn for the plurality of candidate polyadenylation sites.
Preference comparison across sequences to determine antisense effect
Computer processing the plurality of sets of preferences for each of the plurality of genomic sequences with each other to determine an effect of the antisense oligonucleotide on the plurality of candidate polyadenylation sites.
Therapeutic administration that modulates polyadenylation
Administering a therapeutically effective amount of the antisense oligonucleotide, wherein the administered therapeutically effective amount modulates polyadenylation of at least one of the plurality of candidate polyadenylation sites in the subject.
Overall, the claim coverage centers on using a trained algorithm to compute PAS preference outputs from extracted polyadenylation features for candidate sites in reference and antisense-derived variant sequences, comparing the resulting preference sets to determine an antisense effect, and then administering an amount that modulates polyadenylation at candidate sites in the subject.
Stated Advantages
Documented Applications
In-silico therapy simulation by simulating antisense blocking and predicting shifts in polyadenylation site usage.
PAS selection evaluation and PAS prediction use for determining preference changes near polyadenylation sites, including evaluation for ClinVar variant pathogenicity near PAS.
Genome-wide PAS discovery via scanning the genome to locate potential PAS.
Identification of sequence motifs and tissue-specific polyadenylation features using gradient/saliency analyses, to support polyadenylation-targeted therapeutics framework.
Interested in licensing this patent?