Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
Systems and methods are provided for ranking document data retrieved from a data source in response to a search request. A ranking system retrieves document data from documents in the data source that each includes at least one key term that matches a search term in the search request. For each document, a term frequency value is calculated based on a number of occurrences of the key term in the document. Prefix and suffix term rules are used to determine whether a particular occurrence of the key term in a particular document should be included in determining a term weight value for that particular occurrence of the key term. A relevancy ranking value is determined for each document based on the corresponding term frequency and term weight values. The document data is displayed according to each document's corresponding relevancy ranking value.
Core Innovation
The document describes a document ranking system that retrieves documents matching key terms from a data source in response to a search request and ranks the retrieved documents using term frequency and negation-aware term weighting. A term frequency module queries the data source to identify a plurality of documents, each document comprising a key term matching a search term, and determines a term frequency value comprising a total number of occurrences of the key term in each particular document.
A negation module retrieves at least one negation term and at least one negation rule, then compares the negation term according to the negation rule to other terms in each document within a selected proximity of each occurrence of the key term to determine if each occurrence has a negative context. The negation module determines that a negation term matches another term within the selected proximity according to the negation rule and excludes that particular occurrence of the key term in a document for having the negative context.
Based on each occurrence of the key term that has not been excluded, the negation module determines a corresponding term weight value for the key term in each document. The ranking module determines a corresponding relevancy ranking value for each document based on the corresponding term frequency value and corresponding term weight value, and a user interface module generates a list of document data for display identifying the plurality of documents in order based on the corresponding relevancy ranking value.
The ranking system is further characterized by using tokenization/tagging to mark positive context terms and negative context terms, and by providing document analytics that compares predictive documents versus result documents for accuracy based on the term-context tags.
Claims Coverage
Independent claims cover a ranking application and a ranking system/method that use term frequency, a negation module with negation terms and negation rules within a selected proximity, term weighting based on non-excluded occurrences, and a ranking module that computes a relevancy ranking value for ordering display. Across the independent claim set, inventive features include negation-aware exclusion of key-term occurrences and document relevance computations based on term frequency and term weight, including a ratio-based document relevance formulation.
Term frequency from key-term occurrences in documents
Query the data source to identify a plurality of documents that each comprises a key term matching a search term in the search request; determine a corresponding term frequency value for the key term in each document comprising a total number of occurrences of the key term in a particular document.
Negation-aware negative context exclusion within selected proximity
Retrieve at least one negation term and at least one negation rule from a memory; compare the at least one negation term according to the at least one negation rule to other terms in each document within a selected proximity of each occurrence of the key term to determine if each occurrence of the key term has a negative context; determine that the at least one negation term matches another term within the selected proximity according to the at least one negation rule and exclude a particular occurrence of the key term in the document for having the negative context.
Term weight based on non-excluded key-term occurrences
Determine a corresponding term weight value for the key term in each document based on each occurrence of the key term that has not been excluded.
Relevancy ranking value based on term frequency and term weight
Determine a corresponding relevancy ranking value for each document based on the corresponding term frequency value and corresponding term weight value; the user interface module generates a list of document data for display identifying the documents in order based on the corresponding relevancy ranking value.
Prefix/suffix negation rules within same sentence proximity
Retrieve negation terms comprising prefix negation terms and suffix negation terms, and negation rules comprising a prefix negation rule that a prefix negation term appear before the key term in a same sentence as the key term within a selected proximity and a suffix negation rule that a suffix negation term appear after the key term in the same sentence as the key term within the selected proximity; exclude key-term occurrences when the negation term matches another term according to the negation rules within the selected proximity.
Document relevance as ratio of term weight to term frequency
Determine a corresponding relevancy ranking value for each document where the relevancy ranking value comprises a document relevance value comprising a ratio comprising the term weight value in a numerator of the ratio and the term frequency value in a denominator of the ratio.
Term weight computed from negated occurrences
Determine term weight value based on each occurrence of the key term that has not been excluded; determine a corresponding term weight value comprising a difference between term frequency and negated occurrences, where the negation module counts negated occurrences of the key term within the selected proximity.
Document relevance equation involving TW and TF
The ranking module calculates the document relevance value using an equation involving TW and TF.
Ordering display list of ranked documents or records
Generate a list of document data for display identifying each document or record in an order based on the corresponding relevancy ranking value of each document or record.
Across independent claims, the core inventive structure is a ranking application, system, or method that computes term frequency for key-term occurrences, determines negative context using negation terms and negation rules within a selected proximity to exclude certain occurrences, computes term weight from the remaining occurrences, and ranks documents or records for display using a relevancy ranking value based on term frequency and term weight, including ratio-based document relevance and an equation form using TW and TF.
Stated Advantages
Documented Applications
Ranking document data in response to search requests by displaying ranked documents in order of relevancy ranking value.
Performing document analytics by comparing predictive documents versus result documents for accuracy based on term-context tags marked for positive and negative context terms.
Generating lists of ranked document data for display using a user interface module.
Interested in licensing this patent?