awesome-repositories.com
Blog
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectDespreCum realizăm clasamentulPresăServer MCP
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
alibaba avatar

alibaba/DataX

0
View on GitHub↗
17,241 stele·5,658 fork-uri·Java·2 vizualizări

DataX

DataX is a distributed data integration framework and plugin-based ETL tool designed for synchronizing large datasets between heterogeneous sources and destinations. It functions as a JDBC data migration engine and offline synchronization tool, enabling the movement of data between relational databases, NoSQL stores, and object storage.

The system utilizes a plugin-based connector architecture that decouples reader and writer logic, allowing it to map and transform data types across different storage engines using a standardized internal representation. This design supports heterogeneous data pipelines where source-specific data is converted into compatible target types to ensure cross-platform compatibility.

The framework provides comprehensive capabilities for data extraction, including support for columnar formats, incremental synchronization via SQL filtering, and archive decompression. Its writing surface includes batch commit operations, idempotent write strategies to maintain consistency during retries, and the ability to execute pre- and post-synchronization SQL scripts.

Performance is managed through task-level parallelism, throughput control to regulate memory and network traffic, and batch-based write buffering to increase ingestion speed.

Features

  • Data Extraction - Provides core tools for isolating and retrieving specific data points from relational database sources.
  • Large Scale Data Integration Frameworks - Functions as a distributed framework for synchronizing massive volumes of data between heterogeneous sources and destinations.
  • Plugin-Based Architectures - Built on a plugin-based architecture that decouples reader and writer logic to support heterogeneous system integration.
  • Analytical Data Loads from Object Storage - Loads data from cloud object storage into a transportable format for analytical processing.
  • Relational Data Synchronization - Synchronizes relational database records from sources like MySQL, PostgreSQL, and Oracle into target tables via JDBC.
  • Cross-Database Data Migrations - Migrates large datasets between disparate database engines using a standardized internal representation for cross-platform compatibility.
  • Database-Specific Extractions - Enables record extraction from OceanBase databases via JDBC and SQL queries for migration.
  • Data Insertion Interfaces - Implements a programmatic interface for inserting records into target systems via standard JDBC statements.
  • Data Integration Pipelines - Orchestrates the movement and routing of data between object storage, graph databases, and analytical warehouses via a plugin architecture.
  • Intermediate Representations - Employs internal data models that normalize diverse input formats into a consistent structure for uniform processing across different storage engines.
  • Data Type Mappings - Translates source data types into compatible target database-specific column formats.
  • Data Extraction - Implements specialized extraction logic to read records from MySQL databases using JDBC and SQL queries.
  • Data Extraction - Retrieves records from remote Oracle databases using JDBC and generated or custom SQL statements.
  • Distributed Batch Processing - Transfers terabyte-scale datasets using parallel extraction and distributed writes to maximize system throughput.
  • Heterogeneous Data Pipelines - Implements pipelines that transform and map data types across different storage engines to ensure cross-platform compatibility.
  • Heterogeneous Data Synchronization - Enables data migration between diverse storage types such as relational databases and NoSQL stores using a standardized internal format.
  • Incremental Data Synchronization - Synchronizes only new or modified records by filtering data using WHERE clauses based on timestamps or IDs.
  • JDBC Migration Engines - Leverages JDBC drivers to extract and load records across a wide variety of relational database management systems.
  • Plugin-Based ETL Frameworks - Uses a plugin-based connector architecture to decouple reader and writer logic, allowing extensions for new heterogeneous data sources.
  • Pre and Post Load SQL Execution - Executes custom SQL statements immediately before or after synchronization tasks to prepare or finalize data.
  • Relational Database Writers - Inserts data into target relational database tables using JDBC connections and custom drivers.
  • Data Extraction - Reads records from remote SQL Server databases using JDBC connections and SQL SELECT statements.
  • Data Extraction - Extracts data from remote PostgreSQL databases using JDBC connections and SQL select statements.
  • Data Partition Parallelism - Supports task-level parallelism by splitting large datasets into independent chunks for concurrent extraction across threads.
  • Hub-and-Spoke Data Flows - Utilizes a hub-and-spoke data flow to route information from specialized readers through a central engine to specialized writers.
  • Data Component Plugins - Integrates new data sources through the implementation of customizable reader and writer plugins.
  • Parallel Data Pipelines - Splits large datasets into concurrent tasks based on primary keys to increase synchronization speed.
  • Object Selection Filters - Supports isolating specific objects for extraction through name-based filtering and directory traversal.
  • Batch Write Buffering - Groups multiple record writes into a single transaction to increase data ingestion speed and reduce network overhead.
  • Custom SQL Execution - Allows the use of custom SQL queries for data extraction to enable complex operations like multi-table joins.
  • Export Throughput Limiters - Regulates memory usage and network traffic by capping batch sizes and thread counts during data import.
  • Data Synchronization Tools - Provides a tool for bulk data migrations and incremental synchronizations between relational databases and NoSQL stores.
  • Data Transformation Functions - Provides built-in functions and custom rules for masking, completing, and filtering data during the migration process.
  • Unstructured Text Converters - Translates raw string values from object storage into structured types such as decimals and integers.
  • Dirty Data Captures - Provides a dirty-data capture mechanism to intercept and isolate records failing type conversion, ensuring pipeline stability.
  • Graph Database Writers - Synchronizes data into Neo4j using Cypher queries to create nodes and relationships.
  • Graph Database Exporters - Exports data into graph databases by converting source records into vertices and edges.
  • Idempotent Write Strategies - Implements idempotent write strategies that clear target partitions before writing to maintain consistency during failed task retries.
  • Bulk Load Optimizations - Moves terabyte-scale data by leveraging temporary distributed storage before triggering optimized bulk load commands.
  • Columnar Tabular Storage - Extracts data from optimized columnar storage files including Parquet and ORC for large-scale processing.
  • NoSQL Data Writers - Transfers structured data into OTS NoSQL databases with multi-version record support.
  • Parallel Storage Writing - Distributes data across multiple threads to write different sub-files simultaneously, improving overall ingestion throughput.
  • Relational Database Drivers - Integrates various relational database types by registering JDBC drivers and adding necessary driver files.
  • Schema Column Mapping - Selects specific columns for import and rearranges their order to align with the destination schema.
  • SQL Data Retrieval - Implements techniques for filtering and extracting specific data from relational tables using SQL WHERE clauses.
  • Structured Data File Extractors - Extracts text and field names from structured data files such as CSV and TXT using custom delimiters.
  • Tabular Object Storage - Transfers tabular data into object storage using structured formats like Parquet, ORC, and CSV.
  • Throughput Controls - Provides mechanisms to cap transfer rates using concurrency channels or byte limits to prevent target system overloading.
  • Column Projection - Prunes data during export by selecting a specific subset of columns and defining their output order.
  • Throughput Controllers - Regulates memory usage and network traffic by capping concurrency levels and batch sizes during the transfer process.
  • Database Batch Writes - Implements batch-based write buffering to group records into single transactions, reducing network overhead and increasing ingestion speed.
  • Transfer Retry Mechanisms - Automatically retries failed operations at the thread or task level to ensure completion during network instability.
  • Data Exchange and ETL - Data synchronization and exchange tool.

Istoric stele

Graficul istoricului de stele pentru alibaba/dataxGraficul istoricului de stele pentru alibaba/datax

Căutare AI

Explorează mai multe repository-uri excelente

Descrie ce ai nevoie în limbaj simplu — AI-ul sortează mii de proiecte open source selectate în funcție de relevanță.

Start searching with AI

Alternative open-source pentru DataX

Proiecte open-source similare, clasificate după numărul de funcționalități comune cu DataX.
  • dlt-hub/dltAvatar dlt-hub

    dlt-hub/dlt

    5,472Vezi pe GitHub↗

    dlt is a Python data ingestion tool and ETL pipeline framework designed to fetch data from diverse sources and persist it into structured destinations. It functions as a schema inference engine that automatically detects data types and flattens nested JSON structures into relational tables, moving data from sources to lakehouses, warehouses, or vector databases. The project distinguishes itself through AI-powered pipeline generation, using large language models to scaffold extraction code and connectors for REST APIs. It also supports multimodal vector storage and specialized population of ve

    Pythondatadata-engineeringdata-lake
    Vezi pe GitHub↗5,472
  • dimitri/pgloaderAvatar dimitri

    dimitri/pgloader

    6,295Vezi pe GitHub↗

    pgloader is a command-line tool that automates the migration of data and schema from various source databases and file formats into PostgreSQL. It combines schema discovery, parallel data pipelines, and type casting into a single, declarative workflow, using PostgreSQL's COPY protocol for high-throughput bulk loading. The tool distinguishes itself by compiling a dedicated command language into concurrent reader-writer pipelines that handle schema introspection, data transformation, and error-resilient batch processing. It supports migrating entire databases from MySQL, MS SQL, SQLite, and Pos

    Common Lispclozure-clcommon-lispcsv
    Vezi pe GitHub↗6,295
  • apache/flink-cdcAvatar apache

    apache/flink-cdc

    6,430Vezi pe GitHub↗

    This project is a streaming data integration framework that captures real-time database changes and synchronizes them with downstream systems. It operates as a distributed streaming ETL and database synchronizer, reading database logs and snapshots to propagate row-level modifications to target sinks. The system supports declarative data integration, allowing users to define source-to-sink data flows using SQL or YAML configurations. It distinguishes itself by automating schema evolution to maintain synchronization when source structures change and ensuring exactly-once delivery and processin

    Javabatchcdcchange-data-capture
    Vezi pe GitHub↗6,430
  • apache/pinotAvatar apache

    apache/pinot

    6,098Vezi pe GitHub↗

    Pinot is a distributed, columnar analytical database designed for high-concurrency, low-latency query processing. It functions as a real-time OLAP datastore, enabling interactive, user-facing analytics by ingesting and querying massive datasets from both streaming and batch sources. The system architecture relies on a centralized controller for cluster coordination and a distributed segment-based storage model to ensure horizontal scalability. The platform distinguishes itself through a hybrid ingestion pipeline that unifies real-time event streams and historical batch data into a single quer

    Java
    Vezi pe GitHub↗6,098
Vezi toate cele 30 alternative pentru DataX→

Întrebări frecvente

Ce face alibaba/datax?

DataX is a distributed data integration framework and plugin-based ETL tool designed for synchronizing large datasets between heterogeneous sources and destinations. It functions as a JDBC data migration engine and offline synchronization tool, enabling the movement of data between relational databases, NoSQL stores, and object storage.

Care sunt principalele funcționalități ale alibaba/datax?

Principalele funcționalități ale alibaba/datax sunt: Data Extraction, Large Scale Data Integration Frameworks, Plugin-Based Architectures, Analytical Data Loads from Object Storage, Relational Data Synchronization, Cross-Database Data Migrations, Database-Specific Extractions, Data Insertion Interfaces.

Care sunt câteva alternative open-source pentru alibaba/datax?

Alternativele open-source pentru alibaba/datax includ: dlt-hub/dlt — dlt is a Python data ingestion tool and ETL pipeline framework designed to fetch data from diverse sources and persist… dimitri/pgloader — pgloader is a command-line tool that automates the migration of data and schema from various source databases and file… apache/flink-cdc — This project is a streaming data integration framework that captures real-time database changes and synchronizes them… apache/pinot — Pinot is a distributed, columnar analytical database designed for high-concurrency, low-latency query processing. It… bruin-data/ingestr — ingestr is a command-line tool for copying and syncing data between different database engines and third-party… dbgate/dbgate — DbGate is a universal database management tool and SQL client that provides a unified interface for querying and…