Method and system for large scale data curation

Inventors

Bates-Haus, NikolausBeskales, GeorgeBruckner, Daniel MeirIlyas, Ihab F.Pagan, Alexander RichterStonebraker, Michael Ralph

Assignees

Tamr Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-10929348-B2

Patent

Publication Date

2021-02-23

Expiration Date


Abstract

An end-to-end data curation system and the various methods used in linking, matching, and cleaning large-scale data sources. The goal of this system is to provide scalable and efficient record deduplication. The system uses a crowd of experts to train the system. The system operator can optionally provide a set of hints to reduce the number of questions sent to the experts. The system solves the problem of schema mapping and record deduplication in a holistic way by unifying these problems into a unified linkage problem.

Core Innovation

The invention describes a computer implemented method for integrating one or more database storage sources by performing object linkage in computer memory on object pairs from the one or more database storage sources. The method separates the object pairs into linked object pairs and non-linked object pairs, and the candidate object pairs include both schema mappings and record duplicates, so linkage is applied jointly to schema mapping and record deduplication.

The method constructs an initial linkage model configured to generate and classify candidate object pairs. Candidate object pairs are generated and then classified as linked object pairs or non-linked object pairs using the initial linkage model. A subset of the candidate object pairs is selected using stratified sampling to send to data experts for labeling, and an enhanced linkage model is constructed based on the candidate object pairs labeled by data experts.

Using the enhanced linkage model, the candidate object pairs are again classified as linked object pairs or non-linked object pairs. The method then produces a global schema for an integrated data source based on the linked object pairs that are schema mappings, where each cluster of linked objects in schema mapping linked object pairs represents one attribute within the global schema. Finally, the method deduplicates by coalescing linked object pairs that are record duplicates, where each cluster of linked objects in record duplicate linked object pairs represents one record within the integrated data source.

Claims Coverage

The provided independent claim covers an end-to-end in-memory object-linkage based integration that unifies schema mapping and record deduplication while using an expert-labeled, model-refined candidate classification workflow. It includes inventive features for initial and enhanced linkage modeling, stratified expert labeling, and global schema construction with deduplication from linked clusters.

In-memory object linkage on integrated object pairs

A computer implemented method for integrating one or more database storage sources by performing object linkage in computer memory on object pairs from the one or more database storage sources, separating the object pairs into linked object pairs and non-linked object pairs, wherein the candidate object pairs include both schema mappings and record duplicates.

Initial linkage model generating and classifying candidates

Constructing an initial linkage model configured to generate and classify candidate object pairs from the one or more database storage sources, wherein the candidate object pairs include both schema mappings and record duplicates; generating candidate object pairs; classifying the candidate object pairs as linked object pairs or non-linked object pairs using the initial linkage model.

Stratified expert labeling for enhanced linkage model

Selecting a subset of the candidate object pairs using stratified sampling to send to data experts for labeling; constructing an enhanced linkage model based on the candidate object pairs labeled by data experts.

Global schema production from schema-mapping linked clusters

Producing a global schema for an integrated data source based on the linked object pairs that are schema mappings, wherein each cluster of linked objects in schema mapping linked object pairs represents one attribute within the global schema.

Record deduplication by coalescing record-duplicate linked clusters

Deduplicating by coalescing linked object pairs that are record duplicates, wherein each cluster of linked objects in record duplicate linked object pairs represents one record within the integrated data source.

Across the independent claim, the inventive scope centers on unifying schema mapping and record deduplication as an object linkage problem. The claim requires initial and enhanced linkage models with stratified sampling to data experts, followed by producing a global schema from linked schema-mapping clusters and deduplicating by coalescing record-duplicate linked clusters.

Stated Advantages

Not explicitly described in patent.

Documented Applications

Not explicitly described in patent.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.