Systems and methods for rapid gene set enrichment analysis
Inventors
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
Systems and methods for rapid gene set enrichment analysis and applications thereof are described. These systems and methods enable improved prioritization of relevant gene sets while maintaining a relatively lower false positive rate. Additionally, the ability to accelerate enrichment analysis and analyze hundreds of thousands of gene sets is described.
Core Innovation
The invention provides “Cosiner™” systems and methods for rapid quantitative gene set enrichment using a vector space model. Disease and query gene sets, as well as a gene-set corpus, are represented as weighted numeric vectors derived from gene expression data. Similarity between gene sets is computed as a function of an angle between their corresponding numeric vectors, including cosine similarity (cosSim), and can be computed as a distance.
The invention further computes per-gene enrichment statistics using the gene-set numeric vectors, including partial cosine, to enable secondary “enrich by” analysis. This per-gene enrichment supports identifying pathways and drug mechanisms associated with enriched gene sets. The described framework includes a direct enrichment approach and a secondary enrichment workflow that builds on the computed per-gene contributions.
The invention scales to very large gene-set and drug libraries by using sparse matrix linear algebra. The described approach includes non-permutation significance testing with Benjamini–Hochberg adjustment and claims lower false positive rates.
Claims Coverage
The partial content includes three independent claims (clm-00001, clm-00024, clm-00036). Each independent claim uses at least one shared core inventive concept: representing disease and therapy gene sets as weighted numeric vectors and using an angle-based similarity between the vectors to identify therapy candidates, with clm-00036 further emphasizing per-gene enrichment statistics.
Angle-based similarity between disease and therapy gene set vectors
For a first gene set represented as a first numeric vector comprising weighted values and a second gene set represented as a second numeric vector comprising weighted values, determine a measure of similarity between the first gene set and the second gene set as a function of an angle between the first numeric vector and the second numeric vector.
Therapy candidate selection using similarity measures
Identify one or more members of the plurality of candidate therapies as candidates for treatment of the disease using the measures of similarity.
Disease gene set representation as weighted numeric vectors of differentially expressed genes
Identify a first gene set comprising differentially expressed genes of cells indicative of the disease as compared to cells not indicative of the disease, wherein the first gene set is represented as a first numeric vector comprising weighted values corresponding to gene expression data of the differentially expressed genes of the first gene set.
Therapy-associated gene set representation as weighted numeric vectors
For each of a plurality of candidate therapies causing changes in gene expression of cells, identify a second gene set corresponding to each of said candidate therapies, wherein the second gene set is represented as a second numeric vector comprising weighted values corresponding to gene expression data of differentially expressed genes of cells treated with said candidate therapy as compared to cells not treated with said candidate therapy.
Per-gene enrichment statistic computed using disease and therapy vectors
Compute a per-gene enrichment statistic for each gene of a subset of the first gene set using the first numeric vector and the second numeric vector.
Across the independent claims, the claim coverage centers on representing disease and therapy gene sets as weighted numeric vectors and computing an angle-based similarity (including cosine similarity in dependents) to select therapy candidates. The independent system of clm-00036 additionally computes per-gene enrichment statistics for genes in a subset of the disease gene set using the same disease and therapy vectors.
Stated Advantages
Claims lower false positive rates.
Non-permutation significance testing with Benjamini–Hochberg adjustment is included.
Sparse matrix linear algebra enables scaling to very large gene-set and drug libraries.
Enables rapid quantitative gene set enrichment.
Documented Applications
Ranking HER2+ breast cancer therapeutic candidates from LINCS/ChEMBL signatures.
Identification and experimental validation of GSK3 inhibitors for metastatic phenotype inhibition, including invasion assay (MDA-MB-231) outcomes.
Interested in licensing this patent?