Data Quality
ML-Powered Deduplication
★ 4.4
Data Catalog for CI/CD
★ 4.0
pip install dedupepip install grai-clientpip install dedupepip install grai-clientData engineers use Dedupe to clean messy customer or product data where the same entity appears with slightly different names or addresses. Engineers label a small training set via the interactive CLI, Dedupe learns a similarity model, then applies it at scale to cluster duplicate records — outputting canonical entity IDs for use in downstream analysis.
Python data engineers use Grai's Python client to register custom data assets and their relationships in the lineage graph. The GitHub Action integration runs impact analysis on pull requests — when a dbt model or database column changes, Grai reports which downstream Python pipelines, dashboards, or ML features depend on it before the change is merged.
Individual Tool Pages