Method of using clusters to train supervised entity resolution in big data

Inventors

Beskales, George Anwar DanyCattori, Pedro GiesemannBatchelor, Alexandra V.Long, Brian A.Bates-Haus, Nikolaus

Assignees

Tamr Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-12242982-B1

Patent

Publication Date

2025-03-04

Expiration Date


Abstract

Record clustering is performed by learning from verified clusters which are used as the source of training data in a deduplication workflow utilizing supervised machine learning.

Core Innovation

The invention relates to supervised record clustering and record deduplication/entity resolution using current cluster membership, proposed cluster membership, and verified cluster membership. A collection of records is provided such that each record has current and proposed memberships, with some current memberships also being verified cluster memberships. An operator indicates, via an interface, whether a record in a requested subset is a member of a cluster, thereby resulting in revised current cluster membership and revised verified cluster membership for the record.

From the revised current cluster memberships and revised verified cluster memberships, the software code creates inferred match training labels including pairs of records with a match label. These inferred match training labels are then used to train a pair-wise classifier. Using the trained pair-wise classifier and the revised verified cluster memberships, the system generates a new proposed cluster membership for each record, thereby producing a record clustering.

The invention further supports label inference that avoids naive cluster-to-pair conversion bias by generating training labels based on revised cluster memberships and revised verified memberships, and in one workflow creating inferred non-match training labels using a log of previous current cluster memberships. Verified-member constraints are used when creating inferred training labels by examining clusters with at least one verified member. In addition, the system uses an interface-driven membership revision workflow that can include verification mode behavior to produce revised memberships for subsequent classifier training and new proposed cluster membership generation.

Claims Coverage

The independent claims are clm-00001, clm-00008, and clm-00016, covering supervised record clustering via interface-driven membership revision, inferred training-label generation (match and, in another workflow, non-match), training of a pair-wise classifier, and generation of new proposed cluster memberships. Across the independent claims, the main inventive features include operator-driven revised current/verified memberships, inferred match training labels for pair-wise classifier training, inferred non-match training labels using a previous-membership log, and generating new proposed cluster memberships from classifier predictions together with revised verified memberships.

Interface-driven revision of current and verified cluster memberships

Requesting, via an interface, a subset of the collection of records and their respective current, proposed, and verified cluster memberships, and indicating via the interface whether a record in the subset is a member of a cluster to produce a revised current cluster membership and a revised verified cluster membership for the record.

Inferred match training labels from revised memberships

Creating inferred match training labels from the revised current cluster memberships and the revised verified cluster memberships, where the inferred match training labels include pairs of records and each pair has a match label.

Training a pair-wise classifier using inferred match labels

Training a pair-wise classifier using the inferred match training labels.

Generating new proposed cluster memberships using the trained pair-wise classifier and revised verified memberships

Generating, using the trained pair-wise classifier and the revised verified cluster memberships, a new proposed cluster membership for each record in the collection of records, thereby producing a record clustering.

Log of previous current cluster memberships for inferred non-match labeling

Storing, for each record, a log of the previous current cluster memberships, and creating inferred non-match training labels from the revised current cluster memberships, the revised verified cluster memberships, and the log of the previous current cluster memberships.

Inferred non-match training labels from revised memberships and membership log

Creating inferred non-match training labels from the revised current cluster memberships, the revised verified cluster memberships, and the log of the previous current cluster memberships, where the inferred non-match training labels include pairs of records and each pair has a non-match label.

Apparatus with interface, memory, pair-wise classifier, and processor-executed software for inferred match training

An apparatus comprising an interface configured to request and revise membership via the interface, memory configured to store revised current and revised verified cluster memberships, a pair-wise classifier, and software code executing in a processor to create inferred match training labels, train the pair-wise classifier, and generate a new proposed cluster membership for each record using the trained pair-wise classifier and revised verified cluster memberships to produce a record clustering.

The independent claims cover an interface-driven workflow that revises current and verified cluster memberships, then uses revised memberships to create inferred match training labels (and, in a separate workflow, inferred non-match training labels using a log of previous current memberships). The inferred labels are used to train a pair-wise classifier, which is then used to generate new proposed cluster memberships for all records, producing record clustering; the apparatus claim mirrors this architecture with interface, memory, classifier, and processor-executed software.

Stated Advantages

Not explicitly described in patent.

Documented Applications

Not explicitly described in patent.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.