Python library using machine learning to perform deduplication and entity resolution on structured data. Particularly useful for identifying and merging duplicate records.
Data engineers use Dedupe to clean messy customer or product data where the same entity appears with slightly different names or addresses. Engineers label a small training set via the interactive CLI, Dedupe learns a similarity model, then applies it at scale to cluster duplicate records — outputting canonical entity IDs for use in downstream analysis.
Python library using machine learning to perform deduplication and entity resolution on structured data. Particularly useful for identifying and merging duplicate records.
Yes, Dedupe is free to use.
Dedupe is listed under the Data Quality category on Python Data Engineering.
// contains affiliate links
Details
Category
Data Quality →Related
| Tool | Pricing | Rating | |
|---|---|---|---|
GE Great Expectationsfeatured Data Validation & Documentation | Free / Paid | ★ 4.7 | → |
YP Ydata Profiling Automated Data Profiling | Free | ★ 4.6 | → |
PY PyDeequ Data Quality for Big Data | Free | ★ 4.5 | → |