awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

12 repositorios

Awesome GitHub RepositoriesWorkflow Orchestration

Systems for scheduling, managing, and monitoring complex data pipelines.

Explore 12 awesome GitHub repositories matching part of an awesome list · Workflow Orchestration. Refine with filters or upvote what's useful.

Awesome Workflow Orchestration GitHub Repositories

Encuentra los mejores repositorios con IA.Buscaremos los repositorios que mejor coincidan usando IA.
  • apache/airflowAvatar de apache

    apache/airflow

    45,902Ver en GitHub↗

    Airflow is a platform for programmatically authoring, scheduling, and monitoring complex data pipelines. It functions as a workflow automation engine that manages the lifecycle of recurring business processes by executing code-defined task dependencies. By representing workflows as directed acyclic graphs, the system ensures that task execution order and data flow are explicitly defined and reliably maintained across distributed computing environments. The platform distinguishes itself through a highly modular, provider-based architecture that decouples core orchestration logic from external

    Programmatic platform for authoring and scheduling data workflows.

    Pythonairflowapacheapache-airflow
    Ver en GitHub↗45,902
  • kestra-io/kestraAvatar de kestra-io

    kestra-io/kestra

    27,073Ver en GitHub↗

    Kestra is a declarative workflow orchestrator designed to manage complex task dependencies and automated processes through versioned configuration files. It functions as a distributed platform that decouples task scheduling from execution by offloading computational workloads to a fleet of worker nodes. The system uses a reactive, event-driven engine to initiate workflows automatically in response to external signals, webhooks, schedules, or file system changes. The platform distinguishes itself through a modular plugin architecture that allows for the integration of custom tasks and external

    Declarative, event-driven platform for managing complex workflows.

    Javaautomationdata-orchestrationdevops
    Ver en GitHub↗27,073
  • dagster-io/dagsterAvatar de dagster-io

    dagster-io/dagster

    14,974Ver en GitHub↗

    Dagster is a data orchestration platform designed to manage the entire lifecycle of data assets through declarative modeling and version-controlled code. It functions as a workflow engine that treats data assets as first-class primitives, allowing teams to define, schedule, and monitor complex pipelines while maintaining clear visibility into lineage, dependencies, and data quality. The platform distinguishes itself by using a code-as-configuration framework that enables standard software engineering practices, such as unit testing and local mocking, to be applied directly to data workflows.

    Library for building and orchestrating data-intensive applications.

    Pythonanalyticsdagsterdata-engineering
    Ver en GitHub↗14,974
  • rudderlabs/rudder-serverAvatar de rudderlabs

    rudderlabs/rudder-server

    4,437Ver en GitHub↗

    Rudder Server es una plataforma de datos de clientes y una tubería de enrutamiento de eventos diseñada para recopilar, transformar y enrutar datos de eventos de clientes desde diversas fuentes a almacenes de datos y herramientas de negocio. Funciona como un resolutor de identidad de clientes, vinculando identificadores de múltiples fuentes para construir un gráfico de identidad unificado y perfiles de clientes conductuales integrales. El sistema se diferencia por sus capacidades de ETL inverso, que envían segmentos y audiencias de clientes procesados desde almacenes de datos de vuelta a aplicaciones operativas de terceros. También proporciona un plano de datos contenedorizado para despliegues en Kubernetes, permitiendo la gestión de la infraestructura de datos como código. La plataforma cubre una amplia gama de capacidades de gestión de datos, incluyendo la transformación de eventos en tiempo real, la validación de esquemas mediante catálogos de datos y la gobernanza de la privacidad. Estas incluyen herramientas para gestionar el consentimiento del usuario, hacer cumplir la residencia de datos dentro de regiones geográficas específicas y enmascarar información de identificación personal durante el tránsito. La instalación y el despliegue de los componentes del plano de datos se gestionan utilizando gráficos de Helm.

    Customer data platform for collecting and activating warehouse data.

    Gobigquerycdpcustomer-data
    Ver en GitHub↗4,437
  • opendcai/dataflowAvatar de OpenDCAI

    OpenDCAI/DataFlow

    2,926Ver en GitHub↗

    DataFlow is an agent-based workflow orchestrator and data pipeline designed to synthesize, clean, and augment large-scale datasets for training large language models. It functions as a synthetic data generator and text curation tool, utilizing an intelligent assistant to assemble modular processing operators into functional pipelines based on user requirements. The project distinguishes itself through a low-code approach, providing a web-based visual interface for designing and monitoring multi-stage execution flows. It features an operator-based registry system that allows for the integratio

    Platform for automating data preparation and AI pipeline workflows.

    Pythondatadata-agentdata-cleaning
    Ver en GitHub↗2,926
  • dagworks-inc/hamiltonAvatar de dagworks-inc

    dagworks-inc/hamilton

    2,528Ver en GitHub↗

    Apache Hamilton helps data scientists and engineers define testable, modular, self-documenting dataflows, that encode lineage/tracing and metadata. Runs and scales everywhere python does.

    Lightweight library for defining data transformations as DAGs.

    Jupyter Notebook
    Ver en GitHub↗2,528
  • multiwoven/multiwovenAvatar de Multiwoven

    Multiwoven/multiwoven

    1,658Ver en GitHub↗

    🔥🔥🔥 Open source Reverse ETL - alternative to hightouch and census.

    Open-source platform for reverse ETL and data activation.

    Rubybigquerycdpcustomer-data-platform
    Ver en GitHub↗1,658
  • bruin-data/bruinAvatar de bruin-data

    bruin-data/bruin

    1,620Ver en GitHub↗

    Build data pipelines with SQL and Python, ingest data from different sources, add quality checks, and build end-to-end flows.

    CLI tool for end-to-end pipeline management and data quality.

    Goanalyticsbigquerydata-analysis
    Ver en GitHub↗1,620
  • pinterest/pinballAvatar de pinterest

    pinterest/pinball

    1,047Ver en GitHub↗

    Pinball is a scalable workflow manager

    DAG-based workflow manager with support for job output passing.

    JavaScript
    Ver en GitHub↗1,047
  • ralfbecher/orionbelt-semantic-layerAvatar de ralfbecher

    ralfbecher/orionbelt-semantic-layer

    55Ver en GitHub↗

    Open-source Semantic Sidecar for AI, analytics, and governed data systems. Compiles declarative YAML models into optimized SQL, semantic context, KPIs, and DQ rules.

    Semantic sidecar for compiling metrics into optimized SQL.

    Python
    Ver en GitHub↗55
  • getstrm/paceAvatar de getstrm

    getstrm/pace

    39Ver en GitHub↗

    Data policy IN, dynamic view OUT: PACE is the Policy As Code Engine. It helps you to programatically create and apply a data policy to a processing platform like Databricks, Snowflake or BigQuery (or plain 'ol Postgres, even!) with definitions imported from Collibra, Datahub, ODD and the like.

    Framework for enforcing data access and transformation agreements.

    Kotlin
    Ver en GitHub↗39
  • dotflow-io/dotflowAvatar de dotflow-io

    dotflow-io/dotflow

    7Ver en GitHub↗

    🎲 Dotflow turns an idea into flow! — Lightweight Python library for execution pipelines

    Python library for building pipelines with retry and scheduling.

    Python
    Ver en GitHub↗7
  1. Home
  2. Part of an Awesome List
  3. Databases & Data
  4. Workflow Orchestration