awesome-repositories.com
ब्लॉग
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेसMCP सर्वर
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

स्केलेबल डेटा पाइपलाइन फ्रेमवर्क्स

रैंकिंग 30 जून 2026 को अपडेट की गई

For स्केलेबल डेटा पाइपलाइन्स बनाने के लिए फ्रेमवर्क, the strongest matches are apache/beam (Apache Beam is a unified batch and streaming data), pathwaycom/pathway (Pathway is a high-performance framework for building unified batch) and apache/flink (Apache Flink is a distributed stream and batch processing). nathanmarz/storm and apache/spark round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.

हाई-थ्रूपुट डेटा प्रोसेसिंग वर्कफ़्लो बनाने, ऑर्केस्ट्रेट करने और मैनेज करने के लिए डिज़ाइन की गई ओपन-सोर्स लाइब्रेरी और डिस्ट्रिब्यूटेड सिस्टम।

स्केलेबल डेटा पाइपलाइन फ्रेमवर्क्स

AI के साथ बेहतरीन रिपॉजिटरी खोजें।हम AI का उपयोग करके सबसे सटीक रिपॉजिटरी खोजेंगे।
  • apache/beamapache का अवतार

    apache/beam

    8,612GitHub पर देखें↗

    Apache Beam is a distributed data pipeline framework and unified data processing model designed to handle both bounded batch data and unbounded real-time streams. It provides a system for building scalable, data-parallel workflows that operate across compute clusters using a single programming model. The framework utilizes a cross-runner pipeline abstraction that decouples the data processing logic from the underlying execution backend, allowing the same pipeline to run on different distributed compute engines. It supports multi-language pipeline development by translating high-level code fro

    Apache Beam is a unified batch and streaming data processing model with a rich connector ecosystem, fault tolerance, and horizontal scalability across multiple runners, making it a flagship data pipeline framework that directly matches this search.

    JavaDirected Acyclic Graph EnginesStateful Processing BackendsUnified Batch and Stream Processing Engines
    GitHub पर देखें↗8,612
  • pathwaycom/pathwaypathwaycom का अवतार

    pathwaycom/pathway

    62,959GitHub पर देखें↗

    Pathway is a high-performance data processing framework designed for building unified batch and streaming pipelines. It functions as an orchestrator for complex data transformations, utilizing a differential dataflow engine to process updates incrementally. By treating static datasets and continuous event streams with identical logic, the platform ensures exactly-once processing semantics and consistent results across diverse data sources. The framework distinguishes itself through its specialized support for real-time artificial intelligence and retrieval-augmented generation. It features in

    Pathway is a high-performance framework for building unified batch and streaming data pipelines, supporting incremental processing, exactly‑once semantics, and a connector ecosystem—making it a strong fit for scalable, distributed pipeline orchestration.

    PythonExactly-Once Processing SemanticsUnified Batch and Stream Processing EnginesIncremental State Management
    GitHub पर देखें↗62,959
  • apache/flinkapache का अवतार

    apache/flink

    26,086GitHub पर देखें↗

    Apache Flink is a distributed processing engine designed for both high-throughput, low-latency data streams and finite batch workloads. It functions as a stateful stream processor and a SQL stream processing engine, providing a unified runtime to execute relational queries and event-based transformations. The system is distinguished by its ability to manage persistent operator state to ensure exactly-once processing guarantees and consistency during failures. It features specialized capabilities for complex event processing to detect temporal patterns and handles out-of-order events using eve

    Apache Flink is a distributed stream and batch processing engine with built-in state management, exactly-once fault tolerance, and backpressure handling, making it a comprehensive data pipeline framework that directly matches the search for scalable, unified processing with a rich connector ecosystem and DAG-based job execution.

    JavaDirected Acyclic Graph EnginesExactly-Once Processing SemanticsUnified Batch and Stream Processing Engines
    GitHub पर देखें↗26,086
  • nathanmarz/stormnathanmarz का अवतार

    nathanmarz/storm

    8,772GitHub पर देखें↗

    Storm is a distributed stream processing framework and fault-tolerant compute engine designed for executing real-time continuous computations across a cluster of machines. It functions as a stateful stream processor and cluster topology manager, enabling the deployment and monitoring of distributed data flow configurations. The system ensures exactly-once semantics by utilizing transactional state management to guarantee that every message in a data stream is processed exactly one time. It further operates as a distributed RPC system, allowing for the integration of non-native languages throu

    Storm is a distributed stream processing framework that lets you build scalable real-time data pipelines as DAGs of spouts and bolts, fitting the category well, though it focuses on stream-only processing rather than unified stream and batch.

    JavaDirected Acyclic Graph PipelinesExactly-Once Processing SemanticsStateful Processing Backends
    GitHub पर देखें↗8,772
  • apache/sparkapache का अवतार

    apache/spark

    43,467GitHub पर देखें↗

    Apache Spark is a unified distributed data processing engine designed for large-scale data analysis and computation graphs. It functions as a distributed machine learning framework, a graph processing system, a real-time stream processor, and a SQL analytics engine. The system enables the execution of distributed SQL querying, large-scale graph analysis, and real-time stream analytics across clusters of machines. It also provides a scalable environment for implementing machine learning algorithms and predictive model development on massive datasets. The engine incorporates relational query e

    Apache Spark is a leading distributed data processing engine that unifies batch and stream processing, provides a rich connector ecosystem, fault tolerance via lineage, DAG-based workflow orchestration, horizontal scalability, stateful stream processing, and backpressure handling—making it a comprehensive data pipeline framework that precisely fits this search.

    ScalaDirected Acyclic Graph Execution Engines
    GitHub पर देखें↗43,467
  • apache/nifiapache का अवतार

    apache/nifi

    5,976GitHub पर देखें↗

    Apache NiFi is a flow-based programming platform that enables the visual design, monitoring, and management of data pipelines. At its core, it provides a web-based visual dataflow designer where users build directed graphs of processors to route, transform, and mediate data movement between any source and destination without writing custom code. The system records fine-grained data provenance for every data item from ingestion to delivery, supporting audit, debugging, and replay of data lineage. The platform distinguishes itself through a zero-master cluster architecture that distributes proc

    Apache NiFi is a full-featured data pipeline platform with visual DAG design, native clustering for horizontal scalability, built-in backpressure handling, and a broad connector ecosystem, making it a comprehensive solution for building scalable, fault-tolerant data pipelines.

    JavaHorizontal Scaling
    GitHub पर देखें↗5,976
  • apache/airflowapache का अवतार

    apache/airflow

    45,902GitHub पर देखें↗

    Airflow is a platform for programmatically authoring, scheduling, and monitoring complex data pipelines. It functions as a workflow automation engine that manages the lifecycle of recurring business processes by executing code-defined task dependencies. By representing workflows as directed acyclic graphs, the system ensures that task execution order and data flow are explicitly defined and reliably maintained across distributed computing environments. The platform distinguishes itself through a highly modular, provider-based architecture that decouples core orchestration logic from external

    Apache Airflow is a widely-used workflow orchestration platform that schedules and monitors DAG-based data pipelines with distributed execution and a rich provider ecosystem, making it a strong fit for batch and scheduled pipelines, though it does not natively unify streaming and batch or handle backpressure.

    PythonDistributed Processing Engines
    GitHub पर देखें↗45,902
  • linkedin/gobblinlinkedin का अवतार

    linkedin/gobblin

    2,267GitHub पर देखें↗

    A distributed data integration framework that simplifies common aspects of big data integration such as data ingestion, replication, organization and lifecycle management for both streaming and batch data ecosystems.

    Gobblin is a distributed data integration framework purpose-built for streaming and batch ingestion, replication, and lifecycle management at scale, directly matching the need for a scalable pipeline framework with unified stream/batch processing, a connector ecosystem, and fault tolerance.

    JavaBig Data FrameworksData IngestionData Ingestion and Integration
    GitHub पर देखें↗2,267
  • spotify/luigispotify का अवतार

    spotify/luigi

    18,676GitHub पर देखें↗

    Luigi is a Python framework designed for building and managing complex batch data pipelines. It functions as a workflow orchestration engine that organizes tasks into directed acyclic graphs, ensuring that jobs execute in the correct logical order based on their dependencies. By utilizing a centralized scheduler, the system coordinates task execution across distributed environments, tracks global workflow state, and prevents redundant processing by verifying the existence of output targets before triggering any work. The project distinguishes itself through a robust state-tracking mechanism t

    Luigi is a Python framework for building batch data pipelines with DAG orchestration and distributed execution, fitting your search for a scalable pipeline framework — it aligns well though it focuses on batch rather than streaming.

    PythonDirected Acyclic Graph Engines
    GitHub पर देखें↗18,676
  • lyft/flytelyft का अवतार

    lyft/flyte

    7,095GitHub पर देखें↗

    Flyte is a distributed machine learning pipeline manager and MLOps workflow engine. It functions as a Kubernetes-native orchestrator used to coordinate data, models, and compute resources for executing machine learning pipelines and autonomous agents at scale. The platform provides specialized infrastructure for the full machine learning lifecycle, including a dedicated model serving platform to deploy trained models as scalable production-ready inference services. It also enables the coordination and state management of autonomous AI agents. The system manages scalable pipeline execution th

    Flyte is a Kubernetes-native workflow engine specialized for machine learning pipelines, making it a solid data pipeline framework for ML workloads with distributed execution and DAG orchestration, though its focus on ML means it may not offer the unified stream-and-batch model or broad connector ecosystem expected of a general-purpose pipeline framework.

    GoAI Workflow OrchestrationDAG-Based OrchestrationDistributed ML Pipeline Managers
    GitHub पर देखें↗7,095
  • benthosdev/benthosbenthosdev का अवतार

    benthosdev/benthos

    8,681GitHub पर देखें↗

    Benthos is a stream processing engine and data integration pipeline used for routing, transforming, and connecting data streams between diverse sources and sinks. It functions as event routing middleware and a change data capture tool, streaming real-time database modifications as discrete events for downstream processing. The system utilizes a declarative pipeline configuration, where data flow and processing logic are defined in a single static file. It features a specialized domain-specific language for mapping, filtering, and enriching data payloads, allowing for complex transformations w

    Benthos is a stream processing engine and data integration pipeline that lets you define data flows declaratively with a rich connector ecosystem and transformation language, fitting as a data pipeline framework, though its distributed processing and fault-tolerance capabilities are less explicitly emphasized than the visitor may expect.

    GoData Ingestion and IntegrationData Integration PipelinesStream Processing Engines
    GitHub पर देखें↗8,681
  • prefecthq/prefectPrefectHQ का अवतार

    PrefectHQ/prefect

    21,640GitHub पर देखें↗

    Prefect is a workflow orchestration platform designed to define, schedule, and monitor complex data pipelines as Python code. It functions as a container-native engine that wraps individual tasks in isolated environments, ensuring consistent dependencies and resource allocation across diverse infrastructure. By utilizing a state-machine-based orchestration model, the system tracks execution progress through discrete transitions and persistent event logs to maintain reliable and observable task processing. The platform distinguishes itself through a decoupled worker-API architecture, which sep

    Prefect is a workflow orchestration platform for building, scheduling, and monitoring data pipelines in Python, which matches the request for a scalable data pipeline framework, though its focus on orchestration and task management means it does not natively provide unified stream-and-batch processing or backpressure handling.

    PythonData Pipeline OrchestrationWorkflow OrchestrationContainer-Native Infrastructure
    GitHub पर देखें↗21,640
  • dagster-io/dagsterdagster-io का अवतार

    dagster-io/dagster

    14,974GitHub पर देखें↗

    Dagster is a data orchestration platform designed to manage the entire lifecycle of data assets through declarative modeling and version-controlled code. It functions as a workflow engine that treats data assets as first-class primitives, allowing teams to define, schedule, and monitor complex pipelines while maintaining clear visibility into lineage, dependencies, and data quality. The platform distinguishes itself by using a code-as-configuration framework that enables standard software engineering practices, such as unit testing and local mocking, to be applied directly to data workflows.

    Dagster is a data orchestration platform that treats pipelines as modular DAGs with first-class asset management, distributed execution, and fault tolerance, covering most of the needed features for scalable pipeline building—though its unified stream-and-batch support is less central than in some alternatives.

    PythonData Pipeline OrchestrationDeclarative OrchestrationWorkflow Orchestration Engines
    GitHub पर देखें↗14,974
  • quantumblacklabs/kedroquantumblacklabs का अवतार

    quantumblacklabs/kedro

    10,889GitHub पर देखें↗

    Kedro is a data science pipeline framework and production toolbox designed to build reproducible, modular workflows using software engineering best practices. It functions as a data engineering orchestrator and catalog manager, bridging the gap between interactive analysis and maintainable production pipelines. The framework distinguishes itself by using a data catalog to decouple data access from processing logic and providing tools to transition analysis from interactive notebooks into structured workflows. It includes a workflow visualization tool that generates visual maps of data pipelin

    Kedro is a data pipeline framework that structures reproducible, modular workflows with a data catalog and orchestration, so it fits the category, but its focus is on single-machine data science pipelines rather than the distributed processing, streaming, fault tolerance, and horizontal scalability you are looking for.

    PythonData Pipeline OrchestrationProduction Data Science ToolboxesData Access Abstractions
    GitHub पर देखें↗10,889
टॉप 10 की एक नज़र में तुलना करें
रिपॉजिटरीस्टार्सभाषालाइसेंसअंतिम पुश
apache/beam8.6KJavaApache-2.017 जून 2026
pathwaycom/pathway63KPythonNOASSERTION16 जून 2026
apache/flink26.1KJavaApache-2.017 जून 2026
nathanmarz/storm8.8KJavaApache-2.016 अग॰ 2017
apache/spark43.5KScalaApache-2.016 जून 2026
apache/nifi6KJavaapache-2.020 फ़र॰ 2026
apache/airflow45.9KPythonApache-2.023 जून 2026
linkedin/gobblin2.3KJavaApache-2.016 जून 2026
spotify/luigi18.7KPythonapache-2.021 फ़र॰ 2026
lyft/flyte7.1KGoApache-2.017 जून 2026

Related searches

  • स्केलेबल डेटा पाइपलाइन्स बनाने के लिए एक फ्रेमवर्क
  • डेटा पाइपलाइन्स के लिए एक Python फ्रेमवर्क
  • डेटा पाइपलाइन्स के लिए एक वर्कफ़्लो ऑर्केस्ट्रेशन टूल
  • Stream processing engine
  • ML पाइपलाइन्स के लिए ऑर्केस्ट्रेटर
  • पाइपलाइन ऑर्केस्ट्रेशन के लिए एक एम्बेड करने योग्य वर्कफ़्लो इंजन
  • पाइपलाइन्स, ETL/ELT और ऑर्केस्ट्रेशन
  • विशाल डेटा के लिए एक डेटाफ्रेम इंजन