An open-source bulk data loader that helps transfer data between various databases, storages, file formats, and cloud services. Embulk supports parallel processing and plugin-based architecture for extensible data transfer pipelines.
Python data engineers invoke Embulk from Python subprocess calls or Airflow BashOperator tasks — generating the YAML config file programmatically from a Python template, then running `embulk run config.yml`. Embulk's parallel file loading is used for bulk data migrations from legacy systems to modern warehouses where Python-native libraries are too slow.
An open-source bulk data loader that helps transfer data between various databases, storages, file formats, and cloud services. Embulk supports parallel processing and plugin-based architecture for extensible data transfer pipelines.
Yes, Embulk is free to use.
Embulk is listed under the ETL Frameworks category on Python Data Engineering.
Details
Related
| Tool | Pricing | Rating | |
|---|---|---|---|
AT Apache Tez DAG-Based Processing Framework | Free | ★ 4.0 | → |
PR Prestofeatured Distributed SQL Query Engine | Free | ★ 4.5 | → |
AH Apache Hive Data Warehouse on Hadoop | Free | ★ 4.3 | → |