Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
A method includes accessing text, identifying a plurality of terms from the text, determining a plurality of term vectors associated with the identified plurality of terms, and clustering the determined plurality of term vectors into a plurality of clusters, the plurality of clusters comprising a first and a second cluster, the first and second clusters each comprising two or more of the determined term vectors. The method further includes creating a first pseudo-document according to the first cluster, creating a second pseudo-document according to the second cluster, identifying a first set of terms associated with the first cluster using latent semantic analysis (LSA) of the first pseudo-document, identifying a second set of terms associated with the second cluster using LSA of the second pseudo-document, and combining the first and second sets of terms into a list of output terms.
Core Innovation
The invention provides multi-concept latent semantic analysis (LSA) query methods that preserve multiple distinct concepts from query text. The method accesses text, identifies a plurality of terms, determines a plurality of term vectors associated with the identified terms, and calculates a weight of each term vector. The term vectors are clustered into a plurality of clusters, where each cluster is related to a distinct concept of the text.
For each concept-related cluster, the method identifies a set of terms associated with the cluster using LSA, and it determines cluster weights based at least on the weights of the term vectors of each cluster. The method determines output-term proportions from the cluster weights by calculating percentages of a list of output terms that should come from each cluster using ratios of the cluster weights to a sum of the cluster weights. It then selects one or more terms from each LSA-derived cluster term set according to the determined percentages, combines the selected terms, and stores the list of output terms having the distinct concepts of the text.
The invention further addresses term importance by using log-entropy statistics to distinguish important term vectors and unimportant term vectors, and it prevents clustering of important term vectors with unimportant term vectors. Additional refinements include log-entropy mixing when creating the list of output terms, optional cleaning of cluster-derived results toward a query pseudo-document using vector normalization and interpolation, and maintaining comprehensive LSA term spaces using libraries of LSA spaces to mitigate vocabulary limitations.
Claims Coverage
The document provides three independent claims: a computer system, a computer-implemented method, and a non-transitory computer-readable medium. Across the independent claims, the inventive features focus on weighted term-vector clustering into distinct concept clusters, LSA-based identification of per-cluster term sets, and proportion-based selection and combination of output terms to represent multiple distinct concepts.
Weighted term-vector clustering into distinct concept clusters
Clustering the determined plurality of term vectors into a plurality of clusters, the plurality of clusters comprising a first cluster related to a first concept of the text and a second cluster related to a second concept of the text, the first concept being distinct from the second concept, the first and second clusters each comprising two or more of the determined term vectors, the clustering comprising grouping two or more of the determined term vectors together based on the determined weights of the two or more term vectors and a distance between the two or more term vectors
LSA-derived term sets per concept cluster
Identifying, using latent semantic analysis (LSA), a first set of terms associated with the first cluster; identifying, using LSA, a second set of terms associated with the second cluster
Cluster-weight-based output term percentages
Determining a first weight associated with the first cluster and a second weight associated with the second cluster, wherein the first weight is based at least on the weights of the term vectors of the first cluster, and wherein the second weight is based at least on the weights of the term vectors of the second cluster; determining a first percentage of a list of output terms that should come from the first cluster and a second percentage of the list of output terms that should come from the second cluster, the first percentage based on a ratio of the first weight to a sum of the first and second weights, the second percentage based on a ratio of the second weight to the sum of the first and second weights
Proportional term selection and combination into multi-concept output list
Selecting one or more terms from the first set of terms according to the determined first percentage; selecting one or more terms from the second set of terms according to the determined second percentage; combining the selected terms from the first and second sets of terms into the list of output terms, the list of output terms having the first and second concepts of the text; and storing the list of output terms in the one or more memory units
Across the independent claims, the coverage centers on clustering weighted term vectors into distinct concept clusters, extracting concept-specific term sets using LSA, and selecting and combining output terms according to percentages derived from cluster weights so the output list has multiple distinct concepts of the text.
Stated Advantages
Preserves multiple distinct concepts from the query text in the list of output terms.
Documented Applications
Multi-concept LSA query for extracting an output list of terms representing distinct concepts of the text.
Interested in licensing this patent?