Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
Fast record deduplication is accomplished by providing as an input, data records having multiple attributes, and local similarity functions of individual attributes with local similarity thresholds. Bin IDs are then generated based on the local similarity functions and the local similarity thresholds. The Bin IDs are unique identifiers of a respective bin of records, and the bin of records is a set of records that are possibly pairwise similar. Local candidate pairs are identified based on data records that share Bin IDs. The local candidate pairs are aggregated to produce a set of global candidate pairs. The set of global candidate pairs are filtered by deciding whether a pair of data records represents a duplicate.
Core Innovation
The invention provides a computer implemented fast record deduplication approach for data records having multiple attributes. It uses local similarity functions of individual attributes together with local similarity thresholds to generate Bin IDs. The Bin IDs are unique identifiers of bins of records such that all similar record pairs appear in the same bin.
After Bin ID generation, the approach identifies local candidate pairs based on data records that share Bin IDs. The local candidate pairs are then aggregated to produce a set of global candidate pairs. Finally, the global candidate pairs are filtered by deciding whether a pair of data records represents a duplicate.
In a further embodiment reflected in the relevant dependent claims, candidate pairs are identified by forming a Cartesian product of all data records that share the same Bin IDs. The overall workflow supports scalable record deduplication by using attribute-level local similarity functions and local thresholds for candidate generation, followed by duplicate filtering based on the global candidate pairs.
Claims Coverage
The document contains two independent claims (clm-00001 and clm-00003). Across both, the same set of core inventive steps is covered: binning of records using attribute-level local similarity functions and thresholds to ensure similar pairs share Bin IDs, generating local candidate pairs from shared Bin IDs, aggregating to global candidate pairs, and filtering to decide duplicates, with a dependent claim specifying Cartesian-product enumeration for records in the same bins.
Bin ID generation from attribute-level local similarity and thresholds
Generating Bin IDs based on local similarity functions of individual attributes and the local similarity thresholds, where the Bin IDs are unique identifiers of a respective bin of records and all similar record pairs appear in the same bin.
Local candidate pair identification from shared Bin IDs
Identifying local candidate pairs based on data records that share Bin IDs.
Global candidate aggregation from local candidate pairs
Aggregating the local candidate pairs to produce a set of global candidate pairs.
Duplicate filtering by deciding whether a record pair is a duplicate
Filtering the set of global candidate pairs by deciding whether a pair of data records represents a duplicate.
Cartesian-product candidate enumeration from records sharing Bin IDs
Forming candidate pairs as a Cartesian product of all data records that share the same Bin IDs.
The independent claims consistently cover a full deduplication pipeline that assigns Bin IDs using attribute-level local similarity functions and local similarity thresholds so that similar pairs fall into the same bin, then uses shared Bin IDs to generate local candidate pairs, aggregates them into global candidate pairs, and filters them by duplicate decision; dependent refinements specify Cartesian-product enumeration of candidate pairs from records sharing the same Bin IDs.
Stated Advantages
Not explicitly described in patent.
Documented Applications
Not explicitly described in patent.
Interested in licensing this patent?