Automatic entity resolution with rules detection and generation system

Inventors

OSESINA, Olukayode Isaac • Riopka, Taras P.

Assignees

Aware Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-11816078-B2

Patent

Publication Date

2023-11-14

Expiration Date


Abstract

Entity resolution (i.e., record linkage) involves the analysis/discovering of datasets that refer to the same real world entity. Analysis typically involves transformation and comparison of different fields of the dataset followed by the application of often domain/data specific logic for determining datasets that refer to the same real world entity (e.g., person). Consider, a bulk mailing of product catalogs to potential customers. Some individuals may have numerous public records that identify the individual differently. Illustratively, several records associated with Jane Doe at her current home address may exist: one record with her name listed as J. Doe, a second record as Jane H. Doe, a third record as Doe, Jane, and a fourth record as Jan Doe (a misspelling). Conceivably, the bulk mailing could unwittingly send multiple catalogs to Jane Doe at her current address, one for each name variation. The entity resolution process described herein can overcome such problems.

Core Innovation

An entity resolution system generates linkage features and computes linkage feature values for linkage data instances. Linkage features include field pairs and similarity metrics, and the system assigns linkage feature values to create linkage data instances suitable for building a linkage classifier model that outputs whether records refer to the same real-world entity.

The system automatically labels linkage data instances by using existing entities or by using expected error and variation descriptions. Labeled training instances are then used to build and update the linkage classifier model, including the use of positive label and negative label training instances to address negative-sample imbalance.

To improve training, the system uses entropy-based challenging training instance sampling, including selecting samples in a challenging classification region. The system also performs dimensionality reduction to select an efficient set of link features, such as a dimension reducer that reduces the dimensionality of the link feature set.

The system performs probabilistic linkage modeling to determine whether records refer to the same real-world entity, rather than relying solely on deterministic rules. The described architecture includes an entity resolution device and components that generate linkage features, determine similarity, assign link feature values, label training instances, and manage a template for a link feature set.

Claims Coverage

The patent claims cover 4 inventive features.

Linkage feature generation and scoring

Generating linkage features, including field pairs and similarity metrics, and computing linkage feature values for linkage data instances.

Automatic labeling of linkage data instances

Automatically labeling linkage data instances using existing entities or expected error and variation descriptions, and using positive label and negative label training instances.

Entropy-based challenging training instance sampling

Using entropy-based challenging training instance sampling, including selecting samples in a challenging classification region, to address negative-sample imbalance during training.

Dimensionality reduction of link features

Selecting an efficient set of link features using dimensionality reduction, including a dimension reducer that reduces the dimensionality of the link feature set.

The claims cover linkage feature generation and scoring, automatic labeling, entropy-based challenging training instance sampling, and dimensionality reduction of link features.

Stated Advantages

Automatically labels linkage data instances for building a linkage classifier model using existing entities or expected error and variation descriptions.

Handles negative-sample imbalance during training using entropy-based challenging training instance sampling.

Selects an efficient set of link features using dimensionality reduction.

Provides probabilistic linkage modeling to determine whether records refer to the same real-world entity.

Supports identity data integration and use cases involving fuzzy text matching for search and filtering.

Helps highlight potentially fraudulent or erroneous identity information, including in biometric databases.

Documented Applications

Identity data integration.

Fuzzy text matching for search and filtering.

Highlighting potentially fraudulent or erroneous identity information, including biometric databases.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.