Methods and computer program products for clustering records using imperfect rules

Inventors

Beskales, George Anwar DanyBates-Haus, NikolausIlyas, Ihab F.

Assignees

Tamr Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-11948055-B1

Patent

Publication Date

2024-04-02

Expiration Date


Abstract

Record clustering is performed for a collection of records using training rules, training-rule labels, training data created from a sample of pairs of records, a pair-wise classifier, and a clustering algorithm. Record clustering is also performed for a collection of records using prediction rules, prediction-rule labels, a pair-wise classifier, and a clustering algorithm.

Core Innovation

The invention relates to record deduplication and record clustering using a machine learning pair-wise classifier trained with training data derived from multiple, imperfect human-defined rules. A collection of records is paired to produce pairs of records with accompanying signals, and a set of training rules is applied to those pairs to produce training-rule labels that include a match label, a non-match label, or a null label. The plurality of rule outputs is reconciled to produce rule-based training labels for training data used to train the pair-wise classifier.

After training, a clustering algorithm generates record clustering using the pairs of records with accompanying signals and the pair-wise classifier. The workflow addresses conflicting rule labels by reconciling the plurality of training-rule labels on each pair to produce a single rule-based training label. Rule-based label generation can be produced by evaluating the training rules on sampled pairs, including reconciliation across match, non-match, and null labels.

The invention further integrates multiple rule-driven workflows, including a prediction-rule workflow in which prediction rules produce prediction-rule labels and the pair-wise classifier produces match labels with similarity scores or non-match labels with similarity scores, and those outputs are reconciled to produce pairs of records with similarity scores for clustering. Validation rules can classify pairs of records into match, non-match, or null labels and produce validation metrics, and those metrics can be presented in a user interface for an operator, with optional incorporation of operator-provided point training labels for reconciling training-rule labels.

Claims Coverage

The document includes three independent claims. Across the independent claims, the central inventive approach contains four core inventive features: rule-based labeling (match/non-match/null), reconciliation to produce rule-based training labels or similarity-scored pairs, training a pair-wise classifier for the training-claim path, and applying a clustering algorithm to generate record clustering.

Rule-based training label generation from training rules on pairs

Providing a collection of training rules for training a pair-wise classifier, where each training rule takes as input a pair of records from the collection of records with accompanying signals and produces a training-rule label that is one of a match label, a non-match label, or a null label; evaluating the training rules on each pair in a selected sample to produce training-rule labels for producing training data.

Reconciling plurality of training-rule labels into rule-based training labels

Reconciling the plurality of training-rule labels on each pair of records with a plurality of training-rule labels to produce a rule-based training label for each pair of records, thereby producing training data including pairs of records with rule-based training labels.

Training a pair-wise classifier from rule-based training data

Training, using a machine learning training method, a pair-wise classifier using the training data.

Clustering records using the trained pair-wise classifier and pair signals

Generating, using a clustering algorithm, a clustering for the collection of records using the pairs of records with accompanying signals and the pair-wise classifier, thereby producing a record clustering.

Prediction-rule scoring by reconciling prediction-rule labels with classifier similarity scores

Evaluating prediction rules on each pair of records with accompanying signals to produce prediction-rule labels, evaluating the pair-wise classifier on each pair to produce classifier labels with similarity scores, and reconciling the prediction-rule labels and the classifier labels with similarity scores to produce pairs of records with similarity scores.

Clustering records using similarity-scored pairs

Generating, using a clustering algorithm, a clustering for the collection of records using the pairs of records with similarity scores, thereby producing a record clustering.

Overall, the independent claims cover producing match/non-match/null labels from multiple training rules on record pairs, reconciling those rule outputs into rule-based training labels to train a pair-wise classifier, and applying a clustering algorithm to generate record clustering. The independent prediction-rule path instead produces similarity-scored pairs by reconciling prediction-rule outputs with pair-wise classifier similarity scores, followed by clustering using those similarity-scored pairs.

Stated Advantages

Not explicitly described in patent.

Documented Applications

Not explicitly described in patent.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.