Document crawling systems and methods

Inventors

Wagers, Doug R.

Assignees

Softek Illuminate Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-8285703-B1

Patent

Publication Date

2012-10-09

Expiration Date


Abstract

Systems and methods are provided for crawling and indexing documents stored in a data storage system. A crawler system processes multiple jobs that each correspond to crawling documents in the data storage system. Each job includes priority data and crawling instructions. The crawler system stores each job in a priority queue in a sequence based on the priority data. The crawler system assigns each job in the priority queue to a next available processing module for processing based on the stored sequence. Before processing each job, the crawler system determines whether to segment the job into smaller steps based on the corresponding crawling instructions. If the job is segmented, one of smaller steps is processed to crawl a group of the documents in the data storage system. The remaining steps are stored in the priority queue to wait for processing.

Core Innovation

The disclosure relates to a document crawling and indexing system in which a data management application crawls documents in a data storage system by scheduling crawl job modules. A scheduling module retrieves a plurality of job modules from a data store, where each job module includes corresponding crawling instructions and corresponding priority data for crawling documents. The job modules are stored in a priority queue in a sequence based on the corresponding priority data, and an execution module assigns each job module to one of a plurality of processing modules according to the sequence for processing.

Each assigned job module identifies a step for processing based on the corresponding crawling instructions, where the step comprises crawling a group of the documents. The job module processes the step to crawl the group of the documents in the data storage system, and then determines if at least one additional step for processing is required based on the corresponding crawling instructions. When an additional step is required, the at least one additional step comprises crawling another group of the documents.

After processing determines that at least one additional step for processing is required, the job module is rescheduled to the scheduling module for insertion into the priority queue. The disclosure further describes crawling instructions that can segment crawling into additional steps, where a portion of crawling is executed and remaining steps are rescheduled back into the priority queue. The scheduling behavior is motivated to reduce impact on the system being crawled, prevent job starvation, and limit bandwidth and throughput bottlenecks, while also supporting priority determination based on document type, creation/edit age, job size, status data, and recurrence intervals.

Claims Coverage

The independent claims present four inventive features in a job-module based crawling workflow with priority-queue ordering, step-based document-group crawling, iterative rescheduling, and, in one claim, status-data-based step identification.

Priority-queued job-module scheduling and processing assignment

A scheduling module retrieves a plurality of job modules from a data store, the job modules each comprising corresponding crawling instructions and corresponding priority data; a priority queue receives the job modules and stores each job module in a sequence according to the corresponding priority data; and an execution module assigns each job module to one of a plurality of processing modules according to the sequence for processing.

Step-based crawling of document groups from crawling instructions

For each assigned job module, identify a step for processing based on the corresponding crawling instructions, the step comprising crawling a group of the documents, and process the step to crawl the group of the documents in the data storage system.

Iterative rescheduling of job modules for additional steps

Determine if at least one additional step for processing is required based on the corresponding crawling instructions, the at least one additional step comprising crawling another group of the documents; and reschedule the job module to the scheduling module for insertion into the priority queue.

Using status data in step identification

Identify a step for processing based on the corresponding crawling instructions and the corresponding status data, the step comprising crawling a group of the documents, and the status data indicating whether the step has been processed.

Across the independent claims, the core inventive coverage is a closed-loop crawling workflow that retrieves job modules containing crawling instructions and priority data, orders the job modules in a priority queue, assigns job modules to processing modules based on the priority sequence, executes step-based crawling of groups of documents, determines whether additional steps are required, and reschedules job modules back into the priority queue. One independent claim further incorporates status data into step identification.

Stated Advantages

Reduces impact on the system being crawled.

Prevents job starvation.

Limits bandwidth bottlenecks and throughput bottlenecks.

Documented Applications

Crawling and indexing radiological examination documents and healthcare record/document data in a Picture Archiving and Communication System (PACS) environment, including electronic health records (EHRs) and electronic medical records (EMRs).

Crawling and indexing patient records/documents and imaging data in a PACS.

Crawling structured and unstructured (free text) data, including HTML pages/web pages, using tags for image/text documents.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.