Methods and computer program products for clustering records using imperfect rules
Inventors
Beskales, George Anwar Dany • Bates-Haus, Nikolaus • Ilyas, Ihab F.
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
Record clustering is performed for a collection of records using training rules, training-rule labels, training data created from a sample of pairs of records, a pair-wise classifier, and a clustering algorithm. Record clustering is also performed for a collection of records using prediction rules, prediction-rule labels, a pair-wise classifier, and a clustering algorithm.
Core Innovation
The invention relates to record deduplication and record clustering using a machine learning pair-wise classifier trained with training data derived from multiple, imperfect human-defined rules. A collection of records is paired to produce pairs of records with accompanying signals, and a set of training rules is applied to those pairs to produce training-rule labels that include a match label, a non-match label, or a null label. The plurality of rule outputs is reconciled to produce rule-based training labels for training data used to train the pair-wise classifier.
After training, a clustering algorithm generates record clustering using the pairs of records with accompanying signals and the pair-wise classifier. The workflow addresses conflicting rule labels by reconciling the plurality of training-rule labels on each pair to produce a single rule-based training label. Rule-based label generation can be produced by evaluating the training rules on sampled pairs, including reconciliation across match, non-match, and null labels.
The invention further integrates multiple rule-driven workflows, including a prediction-rule workflow in which prediction rules produce prediction-rule labels and the pair-wise classifier produces match labels with similarity scores or non-match labels with similarity scores, and those outputs are reconciled to produce pairs of records with similarity scores for clustering. Validation rules can classify pairs of records into match, non-match, or null labels and produce validation metrics, and those metrics can be presented in a user interface for an operator, with optional incorporation of operator-provided point training labels for reconciling training-rule labels.
Claims Coverage
The document includes three independent claims. Across the independent claims, the central inventive approach contains four core inventive features: rule-based labeling (match/non-match/null), reconciliation to produce rule-based training labels or similarity-scored pairs, training a pair-wise classifier for the training-claim path, and applying a clustering algorithm to generate record clustering.
Rule-based training label generation from training rules on pairs
Providing a collection of training rules for training a pair-wise classifier, where each training rule takes as input a pair of records from the collection of records with accompanying signals and produces a training-rule label that is one of a match label, a non-match label, or a null label; evaluating the training rules on each pair in a selected sample to produce training-rule labels for producing training data.
Reconciling plurality of training-rule labels into rule-based training labels
Reconciling the plurality of training-rule labels on each pair of records with a plurality of training-rule labels to produce a rule-based training label for each pair of records, thereby producing training data including pairs of records with rule-based training labels.
Training a pair-wise classifier from rule-based training data
Training, using a machine learning training method, a pair-wise classifier using the training data.
Clustering records using the trained pair-wise classifier and pair signals
Generating, using a clustering algorithm, a clustering for the collection of records using the pairs of records with accompanying signals and the pair-wise classifier, thereby producing a record clustering.
Prediction-rule scoring by reconciling prediction-rule labels with classifier similarity scores
Evaluating prediction rules on each pair of records with accompanying signals to produce prediction-rule labels, evaluating the pair-wise classifier on each pair to produce classifier labels with similarity scores, and reconciling the prediction-rule labels and the classifier labels with similarity scores to produce pairs of records with similarity scores.
Clustering records using similarity-scored pairs
Generating, using a clustering algorithm, a clustering for the collection of records using the pairs of records with similarity scores, thereby producing a record clustering.
Overall, the independent claims cover producing match/non-match/null labels from multiple training rules on record pairs, reconciling those rule outputs into rule-based training labels to train a pair-wise classifier, and applying a clustering algorithm to generate record clustering. The independent prediction-rule path instead produces similarity-scored pairs by reconciling prediction-rule outputs with pair-wise classifier similarity scores, followed by clustering using those similarity-scored pairs.
Stated Advantages
Not explicitly described in patent.
Documented Applications
Not explicitly described in patent.
Interested in licensing this patent?