Automatic entity resolution with rules detection and generation system
Inventors
OSESINA, Olukayode Isaac • Riopka, Taras P.
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
Entity resolution (i.e., record linkage) involves the analysis/discovering of datasets that refer to the same real world entity. Analysis typically involves transformation and comparison of different fields of the dataset followed by the application of often domain/data specific logic for determining datasets that refer to the same real world entity (e.g., person). Consider, a bulk mailing of product catalogs to potential customers. Some individuals may have numerous public records that identify the individual differently. Illustratively, several records associated with Jane Doe at her current home address may exist: one record with her name listed as J. Doe, a second record as Jane H. Doe, a third record as Doe, Jane, and a fourth record as Jan Doe (a misspelling). Conceivably, the bulk mailing could unwittingly send multiple catalogs to Jane Doe at her current address, one for each name variation. The entity resolution process described herein can overcome such problems.
Core Innovation
An entity resolution system generates linkage features and computes linkage feature values for linkage data instances. Linkage features include field pairs and similarity metrics, and the system assigns linkage feature values to create linkage data instances suitable for building a linkage classifier model that outputs whether records refer to the same real-world entity.
The system automatically labels linkage data instances by using existing entities or by using expected error and variation descriptions. Labeled training instances are then used to build and update the linkage classifier model, including the use of positive label and negative label training instances to address negative-sample imbalance.
To improve training, the system uses entropy-based challenging training instance sampling, including selecting samples in a challenging classification region. The system also performs dimensionality reduction to select an efficient set of link features, such as a dimension reducer that reduces the dimensionality of the link feature set.
The system performs probabilistic linkage modeling to determine whether records refer to the same real-world entity, rather than relying solely on deterministic rules. The described architecture includes an entity resolution device and components that generate linkage features, determine similarity, assign link feature values, label training instances, and manage a template for a link feature set.
Claims Coverage
The patent claims cover 4 inventive features.
Linkage feature generation and scoring
Generating linkage features, including field pairs and similarity metrics, and computing linkage feature values for linkage data instances.
Automatic labeling of linkage data instances
Automatically labeling linkage data instances using existing entities or expected error and variation descriptions, and using positive label and negative label training instances.
Entropy-based challenging training instance sampling
Using entropy-based challenging training instance sampling, including selecting samples in a challenging classification region, to address negative-sample imbalance during training.
Dimensionality reduction of link features
Selecting an efficient set of link features using dimensionality reduction, including a dimension reducer that reduces the dimensionality of the link feature set.
The claims cover linkage feature generation and scoring, automatic labeling, entropy-based challenging training instance sampling, and dimensionality reduction of link features.
Stated Advantages
Automatically labels linkage data instances for building a linkage classifier model using existing entities or expected error and variation descriptions.
Handles negative-sample imbalance during training using entropy-based challenging training instance sampling.
Selects an efficient set of link features using dimensionality reduction.
Provides probabilistic linkage modeling to determine whether records refer to the same real-world entity.
Supports identity data integration and use cases involving fuzzy text matching for search and filtering.
Helps highlight potentially fraudulent or erroneous identity information, including in biometric databases.
Documented Applications
Identity data integration.
Fuzzy text matching for search and filtering.
Highlighting potentially fraudulent or erroneous identity information, including biometric databases.
Interested in licensing this patent?