Method and system for large scale data curation
Inventors
Bates-Haus, Nikolaus • Beskales, George • Bruckner, Daniel Meir • Ilyas, Ihab F. • Pagan, Alexander Richter • Stonebraker, Michael Ralph
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
An end-to-end data curation system and the various methods used in linking, matching, and cleaning large-scale data sources. The goal of this system is to provide scalable and efficient record deduplication. The system uses a crowd of experts to train the system. The system operator can optionally provide a set of hints to reduce the number of questions send to the experts. The system solves the problem of schema mapping and record deduplication a holistic way by unifying these problems into a unified linkage problem.
Core Innovation
The disclosure describes a computer-implemented method and system that integrate one or more database storage sources by performing object linkage in computer memory. The method generates and classifies candidate object pairs from the one or more database storage sources, where the candidate object pairs include both schema mappings and record duplicates, and separates them into linked object pairs and non-linked object pairs.
An initial linkage model is constructed to generate and classify candidate object pairs, and a subset of the candidate object pairs is selected to be sent to data experts for labeling. An enhanced linkage model is constructed based on the candidate object pairs labeled by the data experts, and the candidate object pairs are then classified again using the enhanced linkage model.
The disclosure ties schema integration and deduplication through object linkage decisions. A global schema is produced based on linked object pairs that are schema mappings, where each cluster of linked objects in schema-mapping linked object pairs represents one attribute within the global schema, and deduplicating is performed by coalescing linked object pairs that are record duplicates, where each cluster of linked objects in record-duplicate linked object pairs represents one record within the integrated data source.
Claims Coverage
Independent claim clm-00001 recites a unified object linkage approach with four inventive features: candidate generation and classification for schema mappings and record duplicates, expert labeling to enhance the linkage model, global schema production from linked schema-mapping clusters, and record deduplication from linked record-duplicate clusters.
Unified object linkage for schema mappings and record duplicates
Constructing an initial linkage model configured to generate and classify candidate object pairs from the one or more database storage sources, wherein the candidate object pairs include both schema mappings and record duplicates, and separating the object pairs into linked object pairs and non-linked object pairs.
Expert-labeled feedback to enhance the linkage model
Selecting a subset of the candidate object pairs to send to data experts for labeling and constructing an enhanced linkage model based on the candidate object pairs labeled by data experts for subsequent classification into linked object pairs or non-linked object pairs.
Global schema discovery from linked schema-mapping clusters
Producing a global schema for an integrated data source based on the linked object pairs that are schema mappings, wherein each cluster of linked objects in schema mapping linked object pairs represents one attribute within the global schema.
Record deduplication by coalescing linked record-duplicate clusters
Deduplicating by coalescing linked object pairs that are record duplicates, wherein each cluster of linked objects in record duplicate linked object pairs represents one record within the integrated data source.
The inventive coverage is directed to a unified in-memory object linkage process that treats schema mappings and record duplicates together within candidate generation and classification, uses data-expert labeling to build an enhanced linkage model, and then derives both a global schema and a deduplicated integrated data source from linked clusters.
Stated Advantages
Scalability of the object linkage based data curation approach for large-scale data curation.
Reduced operator expertise requirements through reliance on expert questions/labels and labeling workflows.
Documented Applications
Integrating one or more database storage sources into an integrated data source by producing a global schema and deduplicating records based on object linkage decisions.
Supporting an end-to-end large-scale data curation workflow that unifies schema mapping and record deduplication as a single object linkage problem.
Interested in licensing this patent?