Systems and processes of extracting unstructured data from complex documents

Inventors

BECKER, Willian E.CARNIEL FURLANETTO, Marco AntonioFREEMAN, Daniel H.SILVA, Natanael dos SantosGLANDER, Kevin P.JOSHI, Vedang H.RODRIGUES DIAS, RobertoCAMPONOGARA, Marcos

Assignees

ADP Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-12333240-B2

Patent

Publication Date

2025-06-17

Expiration Date


Abstract

The present disclosure relates generally to data extraction of complex documents and, more particularly, to systems, processes and computer program products configured to automatically extract unstructured data from complex documents and perform table understanding on the extracted data. For example, the method includes: detecting, by the computer system, one or more tables within a digitized document; classifying, by the computer system, the one or more detected tables into at least a first table type; identifying, by the computer system, headers within the first table type; extracting, by the computer system, data within the headers and body cells of the first table type; and mapping, by the computer system, a relationship between the extracted data within the headers and the body cells.

Core Innovation

The invention relates to a method of document extraction for digitized documents that include unstructured tables containing rate information. A computer system receives a digitized document comprising a first table and a second table associated with unstructured layouts and unstructured content, and one or more deep learning multi-modal models detect the unstructured tables and classify each detected unstructured table as a rates table or a non-rates table based on content within the detected tables and bounded boxes with coordinates.

After detection and classification, the method performs layout analysis for the rates table while excluding the non-rates table from layout analysis. The layout analysis identifies header cells and body cells of the rates table using the one or more deep learning multi-modal models based on the first content and the first layout within the bounded boxes, and identifies text in different table cells and groups the text from the different table cells together into at least one header cell in a combined table as a third layout.

The invention extracts rate values from the first content based on the header cells and the body cells of the rates table. It maps the extracted rate values to the third layout using spatial heuristics based on the first layout and the first content within the bounded boxes, using a nearest identified element with the header cells by minimizing geometrical distance according to at least one of a proximity constraint or an alignment constraint, and presents the graphical user interface of the third layout in a structured format including coordinates of extracted data from the digitized document.

Claims Coverage

The document provides three independent claims: clm-00001, clm-00011, and clm-00018. Across the independent claims, the inventive features focus on deep learning multi-modal table detection, rates-versus-non-rates classification for unstructured layouts, bounded-box coordinate determination, header-cell identification constrained by rates-table selection, extraction of rate values, spatial-heuristics mapping using proximity and/or alignment constraints, and GUI presentation of a structured third layout with coordinates.

Deep learning multi-modal table detection and rates versus non-rates classification

Train one or more deep learning multi-modal models using training datasets comprising a plurality unstructured tables as inputs to detect tables with unstructured layouts in digitized documents and classify the detected unstructured tables as rates tables or non-rates tables based on a content of the detected unstructured tables.

Bounded boxes with coordinates for unstructured table layouts

Determine bounded boxes with coordinates for each of the first table and the second table based on identifying the first layout and the second layout that are unstructured layouts.

Rates-table layout analysis with header-cell identification and combined header cell grouping

Responsive to classifying the first table as the rates table and excluding the second table as the non-rates table, perform layout analysis comprising identification of header cells and body cells of the rates table using the one or more deep learning multi-modal models, where identification of header cells comprises identifying text in different table cells and grouping the text from the different table cells together into at least one header cell in a combined table as a third layout.

Rate value extraction and spatial-heuristics mapping using nearest identified element

Extract rate values from the first content based on the header cells and the body cells of the rates table, and map the extracted rate values to the third layout using spatial heuristics based on the first layout and the first content within the bounded boxes using a nearest identified element with the header cells by minimizing geometrical distance according to at least one of a proximity constraint or an alignment constraint.

Graphical user interface configured and structured presentation with coordinates

Configure a graphical user interface with the extracted rate values mapped to the third layout using the spatial heuristics and according to the at least one of the proximity constraint or the alignment constraint, and provide for presentation on a display device the graphical user interface of the third layout in a structured format including coordinates of extracted data from the digitized document.

The independent claims collectively cover a computer system and program product that use deep learning multi-modal models to detect and classify unstructured tables as rates tables or non-rates tables, determine bounded boxes with coordinates, identify header cells and extract rate values for the rates table only, map extracted data to the third layout using spatial heuristics with nearest-element selection constrained by proximity and/or alignment, and present the results in a structured GUI format including coordinates.

Stated Advantages

Documented Applications

No documented applications found

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.