System and method for healthcare document management

Inventors

Burgess, Harlow

Assignees

Rivia Health Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-11450417-B2

Patent

Publication Date

2022-09-20

Expiration Date


Abstract

A medical expense document management system including a document management server is disclosed. The server is configured to receive from a user client device an image of a medical document, convert the image into a plurality of text elements using OCR, and determine a source of the document. The server is also configured to retrieve data detectors from a database, each data detector associated with a data type anticipated to be in the document, each detector having at least one identifier, at least one direction describing a potential relative direction of a text element having a label associated with the detector, and at least one validation criteria. The server is also configured to identify a potential descriptor by comparing each text element with the at least one identifier, determine if the text element pointed to by one of the directions meets the validation criteria, and store the validated text element.

Core Innovation

The invention describes a healthcare document management system in which a document management server receives an image of a medical document from a user client device and converts the image into a plurality of text elements using optical character recognition, where each text element has content and an absolute position in the document. The system determines whether the medical document is a bill or an explanation of benefits by searching the content of each text element for at least one distinguishing string unique to one document type. It identifies all postal addresses by inspecting the content of each text element for a postal address format, validates each postal address, and places each postal address in a standardized postal address format.

The system determines a source of the medical document by comparing identified standardized postal addresses with a list of postal addresses unique to known sources, where the source is one of a healthcare provider and an insurer. If postal addresses are not found on the list, the system determines the source by examining text elements neighboring each postal address that does not match a postal address found on the user record. Based on the document type, the system retrieves a plurality of data detectors from the database, where each data detector includes identifiers, at least one direction describing a potential relative direction of a labeled text element, and at least one validation criteria describing a valid format and/or a valid range.

The system identifies a table within the document by calculating relative positions of neighboring text elements using the absolute positions of text elements and comparing the relative positions of the plurality of text elements. It locates a header for the table by comparing the content of text elements within the table with the identifiers of the plurality of data detectors, and validates header text elements using the validation criteria of the data detector that identified the header. The system associates validated header text elements with validated text elements in a row and a column described by the header, identifies potential descriptors outside the table by comparing text element content with data detector identifiers, and validates text elements pointed to by the data detector directions using the corresponding validation criteria.

For user verification, the system sends to the user client device the content of each text element associated with one data detector, receives verification, stores verified content in a first document record linked to a user record, compares document records for duplicates, pairs bills and explanations of benefits based on a common date, determines discrepancies in patient responsibility, and notifies the user; with user permission it generates and transmits a billing discrepancy notification to a healthcare provider. The claims also include source-ordered identifiers and directions using source history, update of source history according to which identifier and direction matched the most text elements, generation of a list of payments due from bills, natural language query handling with escalation to a human agent, and direct receipt of an external document record from a healthcare provider server or an insurer server.

Claims Coverage

Independent claim set includes two independent claims, with one centered on full extraction plus duplicate/discrepancy/pairing workflows and the other centered on extraction and verified storage with detector learning and optional structural/table identification. Across the independent claims, the core inventive features include OCR-positioned text elements, document-type identification via distinguishing strings, standardized postal address validation and source inference, database-based data detectors with identifiers, directions, and validation criteria ordered and updated using source-specific history, and association of validated content for user verification and storage.

User record creation and OCR positioned text elements

create a user record associated with a user using information received from the user client device; receive from the user client device an image of a medical document; convert the image of the medical document into a plurality of text elements using optical character recognition, each text element having a content and an absolute position in the document.

Document type determination via distinguishing strings

determine if a document type of the medical document is one of a bill and an explanation of benefits by searching the content of each text element for at least one distinguishing string, each distinguishing string being unique to one document type.

Postal address identification, validation, and standardization for source inference

identify all postal addresses in the medical document by inspecting the content of each text element for a postal address format; validate each postal address; place each postal address in a standardized postal address format; determine a source of the medical document by comparing each identified postal address with a list of postal addresses unique to known sources, wherein the source is one of a healthcare provider and an insurer; determine the source for postal addresses not found on the list by examining text elements neighboring each postal address that does not match a postal address found on the user record.

Database data detectors with identifiers, directions, and validation criteria

retrieve a plurality of data detectors from the database based on the document type, each data detector associated with a data type anticipated to be in the document, each data detector comprising at least one identifier that is one of a potential label and a potential format, at least one direction describing a potential relative direction of a text element having a label associated with the data detector, and at least one validation criteria describing one of a valid format and a valid range.

Source-ordered identifiers and directions using source history

for each data detector, order at least one of the identifiers and the directions according to a history stored in the database and associated with the source; update, for each data detector, the history associated with the source according to which identifier and which direction matched the most text elements of the data type described by the data detector in the document.

Relative-position table identification using absolute positions

identify a table within the document by calculating for each text element of the plurality of text elements a relative position of at least one neighboring text element relative to the text element using the absolute position of the text element, and comparing the relative positions of the plurality of text elements; locate a header for the table by comparing the content of the text elements within the table with the identifiers of the plurality of data detectors and then identifying the data type of the matching text elements, the header being one of a row and a column; validate, for each identified text element in the header, at least one text element within the other of a row and a column described by the identified text element in the header with the validation criteria of the data detector that identified the identified text element in the header; associate, for each identified text element in the header, at least one validated text element within the other of the row and the column described by the identified text element in the header with the data detector that identified the identified text element in the header.

Validated descriptor association with direction-based validation

identify a potential descriptor by comparing the content of each text element not part of the table with the at least one identifier of at least one data detector; determine if the text element pointed to by one of the at least one direction of the data detector used to identify the potential descriptor meets the validation criteria of the data detector; associate the validated text element with the data detector.

User verification, verified storage, and linked document records with duplicate and discrepancy handling

send to the user client device, for each text element associated with one data detector of the plurality of data detector, the content of the text element, for verification from the user; receive a verification message from the user client device; store the verified content in a first document record in the database, the first document record being linked to the user record; compare the first document record with records associated with other medical documents linked to the user record; notify the user through the user client device that the medical document is a duplicate upon determination that the medical record already exists; pair the first document record describing one of an explanation of benefits and a bill with a second document record describing the other of an explanation of benefits and a bill, based upon at least a common date; determine if there is a discrepancy between a patient responsibility according to the first document record and a patient responsibility according to the second document record; notify the user through the user client device of the discrepancy; generate and transmit a billing discrepancy notification to a healthcare provider who is the source of one of the bill associated with one of the first document record and the second document record, in response to receipt by the document management server of permission from the user client device; generate a list of payments due by collecting payment details from each bill described by one of a plurality of document records linked to the user record; send the list to the user client device.

Natural-language query with agent escalation

receive a natural language query from the user client device; parse the natural language query; search a database for data associated with the parsed query; send data associated with the parsed query to the user client device; escalate the query to a human agent when the database lacks matching data, where the returned data is partly retrieved from a user record specific to the user.

Direct intake of external document records

be network-coupled to one or more of a healthcare provider server and an insurer server; directly receive an external document record from one of those servers.

Across the independent claims, the document emphasizes a server-based pipeline that uses OCR to create absolute-positioned text elements, classifies document type using distinguishing strings, standardizes and validates postal addresses to infer source, and uses database-configured data detectors with identifiers, relative directions, and validation criteria, ordered and updated by source-specific history, to identify and validate table and non-table descriptors for user verification and storage; in the broader claim scope it further includes duplicate detection, pairing of bills and explanations of benefits by common date, discrepancy determination and notifications, a list of payments due, natural language query handling, and direct receipt of external document records.

Stated Advantages

Not explicitly described in patent.

Documented Applications

Not explicitly described in patent.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.