12 Repos
Systems for scheduling, managing, and monitoring complex data pipelines.
Explore 12 awesome GitHub repositories matching part of an awesome list · Workflow Orchestration. Refine with filters or upvote what's useful.
Airflow is a platform for programmatically authoring, scheduling, and monitoring complex data pipelines. It functions as a workflow automation engine that manages the lifecycle of recurring business processes by executing code-defined task dependencies. By representing workflows as directed acyclic graphs, the system ensures that task execution order and data flow are explicitly defined and reliably maintained across distributed computing environments. The platform distinguishes itself through a highly modular, provider-based architecture that decouples core orchestration logic from external
Programmatic platform for authoring and scheduling data workflows.
Kestra is a declarative workflow orchestrator designed to manage complex task dependencies and automated processes through versioned configuration files. It functions as a distributed platform that decouples task scheduling from execution by offloading computational workloads to a fleet of worker nodes. The system uses a reactive, event-driven engine to initiate workflows automatically in response to external signals, webhooks, schedules, or file system changes. The platform distinguishes itself through a modular plugin architecture that allows for the integration of custom tasks and external
Declarative, event-driven platform for managing complex workflows.
Dagster is a data orchestration platform designed to manage the entire lifecycle of data assets through declarative modeling and version-controlled code. It functions as a workflow engine that treats data assets as first-class primitives, allowing teams to define, schedule, and monitor complex pipelines while maintaining clear visibility into lineage, dependencies, and data quality. The platform distinguishes itself by using a code-as-configuration framework that enables standard software engineering practices, such as unit testing and local mocking, to be applied directly to data workflows.
Library for building and orchestrating data-intensive applications.
Rudder Server ist eine Customer Data Platform (CDP) und Event-Routing-Pipeline, die darauf ausgelegt ist, Kundendaten zu sammeln, zu transformieren und von verschiedenen Quellen an Data Warehouses und Business-Tools weiterzuleiten. Es fungiert als Customer-Identity-Resolver, der Identifikatoren aus mehreren Quellen verknüpft, um einen einheitlichen Identitätsgraphen und umfassende verhaltensbasierte Kundenprofile zu erstellen. Das System zeichnet sich durch Reverse-ETL-Funktionen aus, die verarbeitete Kundensegmente und Zielgruppen aus Data Warehouses zurück in operative Drittanbieteranwendungen pushen. Es bietet zudem eine containerisierte Datenebene für Kubernetes-Deployments, was die Verwaltung der Dateninfrastruktur als Code ermöglicht. Die Plattform deckt eine breite Palette von Datenmanagement-Funktionen ab, einschließlich Echtzeit-Event-Transformation, Schema-Validierung via Datenkatalogen und Privacy-Governance. Dazu gehören Tools zur Verwaltung der Benutzereinwilligung, zur Durchsetzung der Datenresidenz innerhalb spezifischer geografischer Regionen und zur Maskierung personenbezogener Daten während der Übertragung. Installation und Deployment der Datenebenen-Komponenten werden mittels Helm-Charts verwaltet.
Customer data platform for collecting and activating warehouse data.
DataFlow is an agent-based workflow orchestrator and data pipeline designed to synthesize, clean, and augment large-scale datasets for training large language models. It functions as a synthetic data generator and text curation tool, utilizing an intelligent assistant to assemble modular processing operators into functional pipelines based on user requirements. The project distinguishes itself through a low-code approach, providing a web-based visual interface for designing and monitoring multi-stage execution flows. It features an operator-based registry system that allows for the integratio
Platform for automating data preparation and AI pipeline workflows.
Apache Hamilton helps data scientists and engineers define testable, modular, self-documenting dataflows, that encode lineage/tracing and metadata. Runs and scales everywhere python does.
Lightweight library for defining data transformations as DAGs.
🔥🔥🔥 Open source Reverse ETL - alternative to hightouch and census.
Open-source platform for reverse ETL and data activation.
Build data pipelines with SQL and Python, ingest data from different sources, add quality checks, and build end-to-end flows.
CLI tool for end-to-end pipeline management and data quality.
Pinball is a scalable workflow manager
DAG-based workflow manager with support for job output passing.
Open-source Semantic Sidecar for AI, analytics, and governed data systems. Compiles declarative YAML models into optimized SQL, semantic context, KPIs, and DQ rules.
Semantic sidecar for compiling metrics into optimized SQL.
Data policy IN, dynamic view OUT: PACE is the Policy As Code Engine. It helps you to programatically create and apply a data policy to a processing platform like Databricks, Snowflake or BigQuery (or plain 'ol Postgres, even!) with definitions imported from Collibra, Datahub, ODD and the like.
Framework for enforcing data access and transformation agreements.
🎲 Dotflow turns an idea into flow! — Lightweight Python library for execution pipelines
Python library for building pipelines with retry and scheduling.