awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoAcerca deCómo clasificamosPrensaServidor MCP
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
dagster-io avatar

dagster-io/dagster

0
View on GitHub↗
14,974 estrellas·1,986 forks·Python·apache-2.0·22 vistasdagster.io↗

Dagster

Dagster is a data orchestration platform designed to manage the entire lifecycle of data assets through declarative modeling and version-controlled code. It functions as a workflow engine that treats data assets as first-class primitives, allowing teams to define, schedule, and monitor complex pipelines while maintaining clear visibility into lineage, dependencies, and data quality.

The platform distinguishes itself by using a code-as-configuration framework that enables standard software engineering practices, such as unit testing and local mocking, to be applied directly to data workflows. Its architecture is built on a pluggable execution engine that decouples orchestration logic from the underlying compute, allowing tasks to run across diverse cloud-native, serverless, and containerized environments. Furthermore, it supports partition-aware scheduling, which enables incremental processing and efficient management of high-volume datasets.

Beyond core orchestration, the system provides a comprehensive suite of tools for data platform management, including automated quality governance, infrastructure cost optimization, and centralized asset cataloging. It integrates with enterprise identity providers for access control and offers robust observability features, such as streaming logs and visual lineage tracking, to ensure system health and compliance.

The platform supports a variety of deployment models, ranging from self-hosted and hybrid configurations to a fully managed control plane. It includes specialized utilities for migrating legacy pipelines and operationalizing interactive scripts into production-ready components.

Features

  • Data Pipeline Orchestration - Builds and schedules complex data workflows as version-controlled code to ensure reliable execution.
  • Workflow Orchestration Engines - Coordinates distributed tasks and data dependencies across heterogeneous cloud environments and external infrastructure.
  • Declarative Orchestration - Models data pipelines as a graph of versioned assets where the system automatically determines execution order based on dependency requirements.
  • Data Lineage - Provides visual and searchable mapping of data flows to document relationships between source inputs and final downstream assets.
  • Data Asset Lifecycle Management - Models data assets as first-class primitives to manage lineage, dependencies, and quality from ingestion through delivery.
  • Data Asset Modeling - Models data as declarative assets to track lineage, quality, and dependencies throughout the entire lifecycle.
  • Data Pipeline Definitions - Expresses data workflows, resources, and scheduling logic as version-controlled code to enable standard software engineering practices.
  • Configuration-as-Code Frameworks - Expresses entire data workflows and infrastructure resources as version-controlled code to enable standard software engineering practices.
  • Observability Platforms - Monitors the health, performance, and lineage of data workflows and assets through a unified interface.
  • Data Ingestion - Provides modular tools to ingest and transform raw information from various sources into usable assets for downstream analysis.
  • Data Partitioning - Divides large datasets into logical slices to enable incremental processing, targeted re-runs, and efficient management of high-volume workflows.
  • Data Quality Frameworks - Integrates validation checks and automated policies directly into pipelines to ensure data integrity throughout the asset lifecycle.
  • Data Catalogs - Maintains a centralized view of data assets, workflows, and lineage to help teams discover and reuse components.
  • Orchestration Engines - Decouples the orchestration logic from the compute layer to allow tasks to run across diverse environments like containers or serverless.
  • Managed Control Planes - Provides a hosted control plane with enterprise-grade features and insights for teams requiring fully managed infrastructure.
  • Infrastructure as Code - Defines deployment environments and infrastructure configurations using declarative files to integrate with existing workflows.
  • Infrastructure as Code Practices - Manages data pipelines and infrastructure configurations using software engineering practices like testing and CI/CD.
  • Data Access Governance - Centralizes visibility across teams and environments while enforcing fine-grained access controls and maintaining audit-ready lineage for compliance.
  • Automated Quality Workflows - Embeds validation checks and automated testing directly into data pipelines to ensure data integrity.
  • Pipeline Observability Tools - Streams execution logs and performance metrics to external monitoring platforms to maintain unified observability and system health.
  • Data Analysis and Processing - Orchestration for data assets and pipelines.
  • Data Pipelines - Orchestration platform for data development and production execution.
  • Data Pipelines and Orchestration - Orchestration platform for developing and observing data assets.
  • Data Processing - Data orchestrator for ML and ETL workflows.
  • Workflow Orchestration - Library for building and orchestrating data-intensive applications.
  • Data Engineering - Data orchestrator for machine learning and ETL workflows.
  • Data Orchestration - Data orchestrator for managing complex machine learning workflows.
  • Job Schedulers - Orchestration platform for data assets.
  • Scheduling - Orchestrates data pipelines for analytics and machine learning.
  • Workflow Orchestration - Development, production and observation of data assets.
  • Workflow Scheduling - Data orchestrator for machine learning and ETL pipelines.
  • General Purpose Orchestration - Data orchestrator for machine learning, analytics, and ETL pipelines.
  • Workflow Frameworks - API for defining DAGs to build data-intensive applications.
  • Workflow Scheduling - Data orchestrator designed for machine learning, analytics, and ETL pipelines.
  • Data Storage - Connects to cloud object stores, data warehouses, and lakehouse architectures to read, write, and version data assets securely.
  • Event-Driven Data Pipelines - Executes data tasks based on custom schedules or external events to automate pipeline runs in response to real-time data changes.
  • Distributed Computing - Launches and manages code execution across cloud-native, serverless, and containerized infrastructure to scale processing power.
  • Deployment Environments - Supports self-hosted, hybrid, and managed infrastructure deployment models to avoid vendor lock-in.
  • Resource Cost Management - Tracks and attributes resource consumption and compute expenses to specific data assets and pipeline runs.
  • Secret Management - Retrieves and rotates sensitive API keys and configuration parameters from secure vaults automatically during pipeline execution.
  • Execution Metadata - Captures and indexes execution logs, lineage, and asset state in a centralized store to provide unified visibility.
  • Automated Alerting Workflows - Notifies teams immediately when data quality checks fail or budget thresholds are exceeded.
  • Unit Testing - Verifies the logic and reliability of data processing code using unit, integration, and mock testing frameworks before deployment.
  • Isolated Execution Environments - Creates ephemeral, production-mirroring sandboxes for pull requests to validate changes end-to-end.
  • Local Development Tools - Provides isolated interfaces and testing frameworks that allow developers to simulate production environments and validate pipeline logic locally.
  • Production Operationalization - Converts interactive notebooks and shell scripts into production-ready pipeline components to eliminate manual deployment overhead.
  • Cloud Storage - Reads and writes data objects to cloud storage buckets to provide a resilient persistence layer for pipeline execution.
  • Audit Logging - Maintains comprehensive logs of user actions and system changes to ensure compliance and visibility into operational history.
  • Data Workflow Integrations - Incorporates conversational models and lifecycle management tools directly into data workflows to automate complex processing.
  • Local Development Environments - Provides isolated interfaces for building, debugging, and mocking data workflows locally before production deployment.
  • Data Synchronization and Consistency - Ensures reliable and uniform data outputs across different sources and timing intervals to maintain a single source of truth.
  • User Identity Management - Integrates with enterprise identity providers to enforce role-based permissions and automate user provisioning through standard authentication protocols.
  • Legacy Migration Strategies - Converts existing workflow definitions into modern data assets while maintaining backward compatibility during system transitions.

Historial de estrellas

Gráfico del historial de estrellas de dagster-io/dagsterGráfico del historial de estrellas de dagster-io/dagster

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Start searching with AI

Alternativas open-source a Dagster

Proyectos open-source similares, clasificados según cuántas características comparten con Dagster.
  • prefecthq/prefectAvatar de PrefectHQ

    PrefectHQ/prefect

    21,640Ver en GitHub↗

    Prefect is a workflow orchestration platform designed to define, schedule, and monitor complex data pipelines as Python code. It functions as a container-native engine that wraps individual tasks in isolated environments, ensuring consistent dependencies and resource allocation across diverse infrastructure. By utilizing a state-machine-based orchestration model, the system tracks execution progress through discrete transitions and persistent event logs to maintain reliable and observable task processing. The platform distinguishes itself through a decoupled worker-API architecture, which sep

    Pythonautomationdatadata-engineering
    Ver en GitHub↗21,640
  • datahub-project/datahubAvatar de datahub-project

    datahub-project/datahub

    12,141Ver en GitHub↗

    DataHub is a metadata management platform designed to unify technical, operational, and business context across diverse data ecosystems. By utilizing a graph-based metadata model and an event-driven ingestion architecture, it creates a centralized source of truth that maps complex data relationships, lineage, and ownership. This foundational framework enables organizations to maintain a synchronized view of their data landscape, supporting both human-led discovery and automated data operations. The platform distinguishes itself through its focus on grounding artificial intelligence and autono

    Pythondata-catalogdata-discoverydata-governance
    Ver en GitHub↗12,141
  • apache/airflowAvatar de apache

    apache/airflow

    45,902Ver en GitHub↗

    Airflow is a platform for programmatically authoring, scheduling, and monitoring complex data pipelines. It functions as a workflow automation engine that manages the lifecycle of recurring business processes by executing code-defined task dependencies. By representing workflows as directed acyclic graphs, the system ensures that task execution order and data flow are explicitly defined and reliably maintained across distributed computing environments. The platform distinguishes itself through a highly modular, provider-based architecture that decouples core orchestration logic from external

    Pythonairflowapacheapache-airflow
    Ver en GitHub↗45,902
  • apache/incubator-airflowAvatar de apache

    apache/incubator-airflow

    45,840Ver en GitHub↗

    This project is a Python workflow orchestration platform and programmatic data pipeline engine used to author, schedule, and monitor complex data pipelines. It functions as a directed acyclic graph manager and scheduler, allowing users to define data movement and transformation tasks as code to ensure precise execution order and maintainability. The platform distinguishes itself by treating workflows as code, enabling pipelines to be versioned and tested through a standard programming language. It utilizes a system of extensible operators to encapsulate integration logic and employs a templat

    Python
    Ver en GitHub↗45,840
Ver las 30 alternativas a Dagster→

Preguntas frecuentes

¿Qué hace dagster-io/dagster?

Dagster is a data orchestration platform designed to manage the entire lifecycle of data assets through declarative modeling and version-controlled code. It functions as a workflow engine that treats data assets as first-class primitives, allowing teams to define, schedule, and monitor complex pipelines while maintaining clear visibility into lineage, dependencies, and data quality.

¿Cuáles son las características principales de dagster-io/dagster?

Las características principales de dagster-io/dagster son: Data Pipeline Orchestration, Workflow Orchestration Engines, Declarative Orchestration, Data Lineage, Data Asset Lifecycle Management, Data Asset Modeling, Data Pipeline Definitions, Configuration-as-Code Frameworks.

¿Qué alternativas de código abierto existen para dagster-io/dagster?

Las alternativas de código abierto para dagster-io/dagster incluyen: prefecthq/prefect — Prefect is a workflow orchestration platform designed to define, schedule, and monitor complex data pipelines as… datahub-project/datahub — DataHub is a metadata management platform designed to unify technical, operational, and business context across… apache/airflow — Airflow is a platform for programmatically authoring, scheduling, and monitoring complex data pipelines. It functions… apache/incubator-airflow — This project is a Python workflow orchestration platform and programmatic data pipeline engine used to author,… dbt-labs/dbt-core — dbt-core is a command-line framework for transforming data within a warehouse using modular SQL and version control.… spotify/luigi — Luigi is a Python framework designed for building and managing complex batch data pipelines. It functions as a…