System, method and computer readable medium for web crawling
Inventors
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
In a web crawler, a URL selection module selects URLs for pages to be downloaded. The URL selection module accesses an interaction data store that stores interaction data for web pages, including interaction data that indicates human interactions with the pages. To reduce the effects of link farms, the URL selection module filters the URLs to select only those URLs that have human interaction histories and provides the selected URLs to a download module for web page downloading.
Core Innovation
A web crawling procedure selects a uniform resource locator (URL) by querying a URL data store for URLs corresponding to webpages that have not been previously downloaded. In response to the querying, a plurality of URLs corresponding to webpages not previously downloaded during the web crawling procedure is retrieved, and for each particular URL, one or more additional URLs within the URL data store are identified where those additional URLs correspond to webpages that have been downloaded and that contain links to the particular URL.
User interaction event data associated with the identified additional URLs and that include links to the particular URL is retrieved from an interaction data store. From the plurality of retrieved URLs, a highest ranked URL is selected based on the user interaction event data retrieved from the interaction data store associated with the plurality of retrieved URLs, and selecting the highest ranked URL comprises performing source element ranking and attention shift analysis for one or more web pages that include links to the particular URL.
The source element ranking and attention shift analysis comprises detecting user interactions with content areas of each web page that include links to the particular URL, wherein the user interactions comprise clicks, hints, or lingers on the content areas. Based on the detected user interactions with the content areas, the one or more web pages that include links to the particular URL are determined to be valid web pages, the selected URL is submitted to a download module, and the downloaded webpage is processed to extract any URLs within the downloaded webpage and add the extracted URLs to the URL data store.
Claims Coverage
Independent claims are directed to a method, an apparatus, and a non-transitory computer-readable storage medium for selecting a highest ranked URL during web crawling. Across the independent claims, there are at least two core inventive features: (i) ranking and selecting candidate URLs using user interaction event data via source element ranking and attention shift analysis, and (ii) downloading the selected URL and feeding extracted URLs back into the URL data store.
Selecting a highest ranked URL using interaction data and attention shift analysis
Querying a URL data store for webpages not previously downloaded, retrieving a plurality of candidate URLs, identifying additional previously downloaded URLs that contain links to a particular URL, retrieving user interaction event data from an interaction data store for the additional URLs, and selecting a highest ranked URL from the plurality based on the retrieved user interaction event data, wherein selection comprises performing a source element ranking and attention shift analysis.
Valid web page determination using detected user interactions
Detecting user interactions with content areas of one or more web pages that include links to the particular URL, wherein user interactions comprise clicks, hints, or lingers on the content areas, and determining, based on the detected user interactions with the content areas, that the one or more web pages that include links to the particular URL are valid web pages.
Downloading the selected URL and adding extracted URLs back to the URL data store
Submitting the selected URL to a download module, downloading a webpage corresponding to the selected URL, processing the downloaded webpage to extract any URLs within the downloaded webpage, and adding the extracted URLs to the URL data store.
Selecting and ranking using a predetermined human interaction behavior
Selecting a highest ranked URL based on retrieving at least one predetermined human interaction behavior, where the selection comprises performing source element ranking and attention shift analysis using detected user interactions with content areas comprising one or more of clicks, hints, or lingers.
Ranking by out-click user behavior action
Ranking multiple URLs according to at least one user behavior action after clicking out from web pages that include links to the URLs, including selecting a highest ranked URL based on out-click.
Ranking URLs by location within known human attentive areas
Ranking multiple URLs according to the location of each URL within known human attentive areas of a web page that contains links to the URLs and selecting a highest ranked URL based on that location.
Conditioned ranking by content interest scores
Querying a content interest data store to determine whether a downloaded web page includes content interest data and, when content interest data is determined to be present, ranking one or more elements of the downloaded web page by their content interest scores.
The independent claims cover selecting a highest ranked URL during web crawling by querying unseen webpages, retrieving candidate URLs, and ranking candidates using user interaction event data through source element ranking and attention shift analysis, including detecting clicks, hints, or lingers and determining linked pages as valid. The selected URL is downloaded, its URLs are extracted, and the extracted URLs are added to the URL data store. Additional independent-claim refinements cover ranking based on predetermined human interaction behavior, out-click actions, URL location within known human attentive areas, and optional element ranking using content interest scores conditioned on content interest data presence.
Stated Advantages
Reduce link-farm and spam crawls.
Improve efficiency and page ranking quality.
Documented Applications
Web crawling procedures that focus crawling on human-used pages to improve efficiency and page ranking quality.
Link-farm and spam filtering in web crawling by selecting URLs based on human user interaction data.
Interested in licensing this patent?