Method of using clusters to train supervised entity resolution in big data
Inventors
Beskales, George Anwar Dany • Cattori, Pedro Giesemann • Batchelor, Alexandra V. • Long, Brian A. • Bates-Haus, Nikolaus
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
Record clustering is performed by learning from verified clusters which are used as the source of training data in a deduplication workflow utilizing supervised machine learning.
Core Innovation
The invention is a method of record clustering that combines user verification with supervised machine learning. It provides a collection of records, each record having a current cluster membership, a proposed cluster membership, and possibly a verified cluster membership. A subset of the collection is displayed on a user interface with each record’s current, proposed, and verified cluster memberships for user review.
The user indicates, via the user interface display, for a record in the displayed subset whether the record is a member of a cluster. This results in a revised current cluster membership and a revised verified cluster membership for the record, and the method stores the revised memberships in memory. The method also stores a log of previous current cluster memberships and uses that information to create inferred training labels, including pairs of records labeled with match or non-match.
The method trains a pair-wise classifier using the inferred training labels and then generates a new proposed cluster membership for each record in the collection using the classifier and the revised verified cluster memberships. This produces the record clustering, with the revised verified cluster memberships incorporated into the process that assigns new proposed memberships to the records.
Claims Coverage
The partial content identifies one independent claim (clm-00001) and multiple dependent claims that refine specific steps and constraints. The independent claim covers an end-to-end record clustering method driven by a user interface, revised verified memberships, inferred match/non-match training labels, a pair-wise classifier, and generation of new proposed memberships for producing record clustering.
User interface-driven revision of cluster memberships
Displaying a subset of records with current, proposed, and verified cluster memberships, receiving user input indicating membership for displayed records, and resulting in revised current cluster membership and revised verified cluster membership for each such record.
Record revision history and membership log storage
Storing, in memory, for each record the revised current cluster membership and revised verified cluster membership, and storing a log of the previous current cluster memberships.
Inferred training labels from revised memberships and previous membership log
Creating inferred training labels from the revised current cluster memberships, revised verified cluster memberships, and the log of previous current cluster memberships, where inferred training labels include pairs of records labeled match or non-match.
Training a pair-wise classifier from inferred training labels
Training a pair-wise classifier using the inferred training labels.
Classifier-driven proposed membership generation using revised verified cluster memberships
Generating, using the pair-wise classifier and the revised verified cluster memberships, a new proposed cluster membership for each record in the collection to produce the record clustering.
Pair generation, match/non-match prediction, and clustering in step (g)
Transforming records into record pairs, using the pair-wise classifier to predict match versus non-match for each pair, and running a clustering algorithm to generate proposed cluster membership for each record using the predicted pairs and the revised verified cluster memberships.
Cluster-based construction of inferred match-labeled training pairs from verified-member clusters
In creating inferred training labels, examining clusters with at least one verified member, selecting records from such clusters, creating match-labeled record pairs from the selected records, and selecting a sample of the formed pairs.
Cluster-based construction of inferred non-match-labeled training pairs using membership differences and membership log
In creating inferred training labels, examining clusters with at least one verified member, selecting records with differing prior and current cluster memberships, selecting corresponding records from the cluster-membership log, forming non-match-labeled pairs, and selecting a sample of the formed pairs.
Conditional logging of previous current memberships when membership changes
Storing a log of the previous current cluster memberships only for records whose revised current cluster membership differs from their previous current cluster membership.
Verification mode controlling how verified cluster membership is revised and used
Allowing user indication of whether verification mode is locked, suggested, or movable, resulting in revised verified cluster membership for records and affecting subsequent generation of new proposed cluster memberships using the revised verified cluster memberships.
Across the identified independent and dependent claim content, the method centers on a user interface that revises current and verified cluster memberships, stores revised memberships and a previous membership log, creates inferred match/non-match training labels from revised memberships and prior membership history, trains a pair-wise classifier, and generates new proposed cluster memberships for all records by incorporating the revised verified cluster memberships.
Stated Advantages
Documented Applications
No documented applications found
Interested in licensing this patent?