1 个仓库
Techniques for moving multi-terabyte datasets using intermediate distributed storage and bulk load commands.
Distinct from Large-Scale Dataset Management: Focuses on the movement and loading process for massive sets, rather than the general management of the stored data.
Explore 1 awesome GitHub repository matching data & databases · Bulk Load Optimizations. Refine with filters or upvote what's useful.
DataX is a distributed data integration framework and plugin-based ETL tool designed for synchronizing large datasets between heterogeneous sources and destinations. It functions as a JDBC data migration engine and offline synchronization tool, enabling the movement of data between relational databases, NoSQL stores, and object storage. The system utilizes a plugin-based connector architecture that decouples reader and writer logic, allowing it to map and transform data types across different storage engines using a standardized internal representation. This design supports heterogeneous data
Moves terabyte-scale data by leveraging temporary distributed storage before triggering optimized bulk load commands.