Review and curation of record clustering changes at large scale

Inventors

WEBBER, Timothy KwokBeskales, George Anwar DanyCunningham, DennisWAGNER RODRIGUEZ, Alan BenjaminCleary, Liam

Assignees

Tamr Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-11321359-B2

Patent

Publication Date

2022-05-03

Expiration Date


Abstract

Methods are provided to represent proposed changes to clusterings for ease of review, as well as tools to help subject matter experts identify clusters that warrant review versus those that do not. These tools make overall assessment of proposed clustering changes and targeted curation practical at large scale. Use of these tools and method enables efficient data management operations when dealing with extreme scale, such as where entity resolution involves clusterings created from data sources involving millions of entities.

Core Innovation

The invention manages large scale data records for entity resolution by comparing a current published clustering dataset to a subsequent proposed clustering dataset within a computer-implemented software program. The current published clustering includes published clusters, where each cluster refers to the same entity, and the proposed clustering includes proposed clusters that differ by data records being added, removed, or modified relative to the current published clustering.

Clusters within the proposed clustering are matched to clusters within the current published clustering, and differences are identified between the current published clustering and the proposed clustering on both cluster and record levels using the matched clusters. The identified differences include added data records, omitted data records, and modified data records where the modifications are data changes to fields of the data records. The differences support review of the proposed clustering after cluster matching and difference identification between the first and second datasets.

The method approves or rejects the proposed clustering based upon a review of the identified differences. Upon approval of the proposed clustering, a new published clustering is created using the proposed clustering, and upon rejection, a new proposed clustering is received for subsequent review. The workflow iteratively moves from a current published clustering to a subsequent proposed clustering through review, difference identification, and an approval or rejection decision.

Claims Coverage

The document provides one independent claim (clm-00001) that defines the core clustering management workflow, including cluster matching, cluster-and-record-level difference identification, and an approval/rejection loop that produces a new published clustering upon approval.

Comparing current published and subsequent proposed entity-resolution clustering

providing a current published clustering dataset defining published clusters and receiving a proposed clustering dataset defining proposed clusters that differ by including records not present, excluding records present, and/or including modified records with data changes to fields.

Matching clusters between proposed and published clustering datasets

matching clusters within the proposed clustering to clusters within the current published clustering.

Identifying differences on cluster and record levels using matched clusters

identifying differences between the current published clustering and the proposed clustering on both cluster and record levels using the matched clusters, including added records, omitted records, and modified records with data changes to fields.

Reviewing differences to approve or reject proposed clustering and iterating

approving or rejecting the proposed clustering based upon a review of the identified differences, creating a new published clustering using the proposed clustering upon approval, and receiving a new proposed clustering for subsequent review upon rejection.

Independent claim clm-00001 covers a full entity-resolution clustering management workflow in a computer-implemented software program: providing a current published clustering, receiving a subsequent proposed clustering, matching clusters across versions, identifying differences at cluster and record levels, and using those differences for review-driven approval or rejection with iterative resubmission until an approved proposed clustering becomes the new published clustering.

Stated Advantages

Documented Applications

Not explicitly described in patent.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.