awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目关于排名机制媒体报道MCP 服务器
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

数据集成与 ETL 平台

排名更新于 2026年6月30日

For 用于数据迁移的连接器, the strongest matches are n8n-io/n8n (n8n is a self-hostable workflow automation platform with over), dagster-io/dagster (Dagster is a data orchestration platform that manages the) and apache/seatunnel (SeaTunnel is an open-source distributed data integration engine that). apache/nifi and streamsets/datacollector round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.

这些开源工具利用预构建的连接器,促进了不同数据源与目标之间的数据无缝迁移。

数据集成与 ETL 平台

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • n8n-io/n8nn8n-io 的头像

    n8n-io/n8n

    192,772在 GitHub 上查看↗

    n8n is a workflow automation platform that combines a visual interface with code-based extensibility to design, orchestrate, and manage automated processes. It provides a comprehensive suite of tools for data transformation, filtering, and storage, allowing users to build complex logic through conditional branching, looping, and sub-workflow execution. The platform supports both pre-built integration nodes and custom code execution in JavaScript or Python, enabling connectivity with a wide range of external services and APIs. The platform includes a suite of generative AI capabilities, such a

    n8n is a self-hostable workflow automation platform with over 400 pre-built connectors, visual and code-based data transformation, scheduling, and monitoring, making it a comprehensive fit for a data integration / ETL platform that moves data between arbitrary sources and destinations.

    TypeScriptBuilt-in Integration Nodes
    在 GitHub 上查看↗192,772
  • dagster-io/dagsterdagster-io 的头像

    dagster-io/dagster

    14,974在 GitHub 上查看↗

    Dagster is a data orchestration platform designed to manage the entire lifecycle of data assets through declarative modeling and version-controlled code. It functions as a workflow engine that treats data assets as first-class primitives, allowing teams to define, schedule, and monitor complex pipelines while maintaining clear visibility into lineage, dependencies, and data quality. The platform distinguishes itself by using a code-as-configuration framework that enables standard software engineering practices, such as unit testing and local mocking, to be applied directly to data workflows.

    Dagster is a data orchestration platform that manages the full lifecycle of data pipelines with scheduling, monitoring, and transformation—making it a solid fit as an open-source data integration tool, though its strength lies in orchestration rather than a vast library of pre-built connectors.

    PythonData Pipeline OrchestrationDeclarative OrchestrationWorkflow Orchestration Engines
    在 GitHub 上查看↗14,974
  • apache/seatunnelapache 的头像

    apache/seatunnel

    9,427在 GitHub 上查看↗

    SeaTunnel is a distributed data integration engine designed to synchronize structured and unstructured data across diverse sources and sinks. It functions as a multi-engine execution framework that can run data integration tasks across different distributed computing backends to optimize workload performance. The project is distinguished by a visual data pipeline designer for configuring workflows without manual code and a specialized change data capture tool for streaming incremental database updates. It also includes an enrichment pipeline that integrates large language models and embedding

    SeaTunnel is an open-source distributed data integration engine that supports batch and real-time sync via connectors to diverse sources and sinks, includes a visual pipeline designer, change data capture, and job execution, making it a strong fit for a self-hostable ETL platform that covers scheduling, transformation, and monitoring.

    JavaBackend-Agnostic Execution LayersDistributed Data EnginesCDC Synchronization
    在 GitHub 上查看↗9,427
  • apache/nifiapache 的头像

    apache/nifi

    5,976在 GitHub 上查看↗

    Apache NiFi is a flow-based programming platform that enables the visual design, monitoring, and management of data pipelines. At its core, it provides a web-based visual dataflow designer where users build directed graphs of processors to route, transform, and mediate data movement between any source and destination without writing custom code. The system records fine-grained data provenance for every data item from ingestion to delivery, supporting audit, debugging, and replay of data lineage. The platform distinguishes itself through a zero-master cluster architecture that distributes proc

    NiFi is a mature open-source data integration platform with a visual flow-based designer, hundreds of processors for arbitrary sources and destinations, built-in transformation, scheduling, data provenance monitoring, and support for both batch and real-time data movement — exactly the comprehensive self-hostable ETL platform described.

    JavaData Pipeline OrchestrationData Pipeline OrchestratorsProcessor Graph Dataflow Models
    在 GitHub 上查看↗5,976
  • streamsets/datacollectorS

    streamsets/datacollector

    0在 GitHub 上查看↗

    StreamSets Data Collector is an open-source data integration platform that provides a wide range of connectors, supports both batch and real-time data movement, includes built-in transformations, scheduling, and monitoring, and can be self-hosted, making it a comprehensive match for your search.

    Data IngestionData Ingestion Pipelines
    在 GitHub 上查看↗0
  • dlt-hub/dltdlt-hub 的头像

    dlt-hub/dlt

    5,472在 GitHub 上查看↗

    dlt is a Python data ingestion tool and ETL pipeline framework designed to fetch data from diverse sources and persist it into structured destinations. It functions as a schema inference engine that automatically detects data types and flattens nested JSON structures into relational tables, moving data from sources to lakehouses, warehouses, or vector databases. The project distinguishes itself through AI-powered pipeline generation, using large language models to scaffold extraction code and connectors for REST APIs. It also supports multimodal vector storage and specialized population of ve

    dlt is an open-source ETL framework that moves data from diverse sources to structured destinations using connectors and schema inference, fitting the core need for a data integration tool though it is more of a library than a full platform.

    PythonData Destination Connectors
    在 GitHub 上查看↗5,472
  • apache/camelapache 的头像

    apache/camel

    6,247在 GitHub 上查看↗

    Apache Camel is an enterprise integration framework and Java integration engine designed to route and mediate data between disparate systems. It functions as a multi-runtime middleware that implements standardized enterprise integration patterns to manage how messages are routed, transformed, and processed. The framework includes a specialized gateway to connect large language models to enterprise data and internal systems using dedicated communication protocols. It utilizes a vast library of pre-built connectors to bridge different communication protocols and enable data exchange between inc

    Apache Camel is a mature integration framework with hundreds of pre-built connectors and built-in routing, transformation, and support for both batch and real-time data movement — it exactly fits the need for an open-source data integration platform, though it is code-driven rather than a graphical platform.

    JavaEnterprise Integration FrameworksEnterprise Integration SuitesData Connector Libraries
    在 GitHub 上查看↗6,247

Related searches

  • 工作流集成平台
  • 自托管的 Fivetran 替代方案
  • 反向 ETL 同步工具
  • 支持所有数据库的通用客户端
  • 数据库间的实时复制
  • 数据流水线工作流编排工具
  • 开源的工作流自动化平台
  • 用于处理关联数据的图数据库