awesome-repositories.com
Blog
MCP
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetServeur MCPÀ proposNotre méthodologiePresse
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
dagster-io avatar

dagster-io/dagster

0
View on GitHub↗
14,974 stars·1,986 forks·Python·apache-2.0·32 vuesdagster.io↗

Dagster

Dagster is a data orchestration platform designed to manage the entire lifecycle of data assets through declarative modeling and version-controlled code. It functions as a workflow engine that treats data assets as first-class primitives, allowing teams to define, schedule, and monitor complex pipelines while maintaining clear visibility into lineage, dependencies, and data quality.

The platform distinguishes itself by using a code-as-configuration framework that enables standard software engineering practices, such as unit testing and local mocking, to be applied directly to data workflows. Its architecture is built on a pluggable execution engine that decouples orchestration logic from the underlying compute, allowing tasks to run across diverse cloud-native, serverless, and containerized environments. Furthermore, it supports partition-aware scheduling, which enables incremental processing and efficient management of high-volume datasets.

Beyond core orchestration, the system provides a comprehensive suite of tools for data platform management, including automated quality governance, infrastructure cost optimization, and centralized asset cataloging. It integrates with enterprise identity providers for access control and offers robust observability features, such as streaming logs and visual lineage tracking, to ensure system health and compliance.

The platform supports a variety of deployment models, ranging from self-hosted and hybrid configurations to a fully managed control plane. It includes specialized utilities for migrating legacy pipelines and operationalizing interactive scripts into production-ready components.

Features

  • Data Pipeline Orchestration - Builds and schedules complex data workflows as version-controlled code to ensure reliable execution.
  • Workflow Orchestration Engines - Coordinates distributed tasks and data dependencies across heterogeneous cloud environments and external infrastructure.
  • Declarative Orchestration - Models data pipelines as a graph of versioned assets where the system automatically determines execution order based on dependency requirements.
  • Data Lineage - Provides visual and searchable mapping of data flows to document relationships between source inputs and final downstream assets.
  • Data Asset Lifecycle Management - Models data assets as first-class primitives to manage lineage, dependencies, and quality from ingestion through delivery.
  • Data Asset Modeling - Models data as declarative assets to track lineage, quality, and dependencies throughout the entire lifecycle.
  • Data Pipeline Definitions - Expresses data workflows, resources, and scheduling logic as version-controlled code to enable standard software engineering practices.
  • Configuration-as-Code Frameworks - Expresses entire data workflows and infrastructure resources as version-controlled code to enable standard software engineering practices.
  • Observability Platforms - Monitors the health, performance, and lineage of data workflows and assets through a unified interface.
  • Data Ingestion - Provides modular tools to ingest and transform raw information from various sources into usable assets for downstream analysis.
  • Data Partitioning - Divides large datasets into logical slices to enable incremental processing, targeted re-runs, and efficient management of high-volume workflows.
  • Data Quality Frameworks - Integrates validation checks and automated policies directly into pipelines to ensure data integrity throughout the asset lifecycle.
  • Data Catalogs - Maintains a centralized view of data assets, workflows, and lineage to help teams discover and reuse components.
  • Orchestration Engines - Decouples the orchestration logic from the compute layer to allow tasks to run across diverse environments like containers or serverless.
  • Managed Control Planes - Provides a hosted control plane with enterprise-grade features and insights for teams requiring fully managed infrastructure.
  • Infrastructure as Code - Defines deployment environments and infrastructure configurations using declarative files to integrate with existing workflows.
  • Infrastructure as Code Practices - Manages data pipelines and infrastructure configurations using software engineering practices like testing and CI/CD.
  • Data Access Governance - Centralizes visibility across teams and environments while enforcing fine-grained access controls and maintaining audit-ready lineage for compliance.
  • Automated Quality Workflows - Embeds validation checks and automated testing directly into data pipelines to ensure data integrity.
  • Pipeline Observability Tools - Streams execution logs and performance metrics to external monitoring platforms to maintain unified observability and system health.
  • Data Analysis and Processing - Orchestration for data assets and pipelines.
  • Data Pipelines - Orchestration platform for data development and production execution.
  • Data Pipelines and Orchestration - Orchestration platform for developing and observing data assets.
  • Data Processing - Data orchestrator for ML and ETL workflows.
  • Workflow Orchestration - Library for building and orchestrating data-intensive applications.
  • Data Engineering - Data orchestrator for machine learning and ETL workflows.
  • Data Orchestration - Data orchestrator for managing complex machine learning workflows.
  • Job Schedulers - Orchestration platform for data assets.
  • Scheduling - Orchestrates data pipelines for analytics and machine learning.
  • Workflow Orchestration - Development, production and observation of data assets.
  • Workflow Scheduling - Data orchestrator for machine learning and ETL pipelines.
  • General Purpose Orchestration - Data orchestrator for machine learning, analytics, and ETL pipelines.
  • Workflow Frameworks - API for defining DAGs to build data-intensive applications.
  • Workflow Scheduling - Data orchestrator designed for machine learning, analytics, and ETL pipelines.
  • Data Storage - Connects to cloud object stores, data warehouses, and lakehouse architectures to read, write, and version data assets securely.
  • Event-Driven Data Pipelines - Executes data tasks based on custom schedules or external events to automate pipeline runs in response to real-time data changes.
  • Distributed Computing - Launches and manages code execution across cloud-native, serverless, and containerized infrastructure to scale processing power.
  • Deployment Environments - Supports self-hosted, hybrid, and managed infrastructure deployment models to avoid vendor lock-in.
  • Resource Cost Management - Tracks and attributes resource consumption and compute expenses to specific data assets and pipeline runs.
  • Secret Management - Retrieves and rotates sensitive API keys and configuration parameters from secure vaults automatically during pipeline execution.
  • Execution Metadata - Captures and indexes execution logs, lineage, and asset state in a centralized store to provide unified visibility.
  • Automated Alerting Workflows - Notifies teams immediately when data quality checks fail or budget thresholds are exceeded.
  • Unit Testing - Verifies the logic and reliability of data processing code using unit, integration, and mock testing frameworks before deployment.
  • Isolated Execution Environments - Creates ephemeral, production-mirroring sandboxes for pull requests to validate changes end-to-end.
  • Local Development Tools - Provides isolated interfaces and testing frameworks that allow developers to simulate production environments and validate pipeline logic locally.
  • Production Operationalization - Converts interactive notebooks and shell scripts into production-ready pipeline components to eliminate manual deployment overhead.
  • Cloud Storage - Reads and writes data objects to cloud storage buckets to provide a resilient persistence layer for pipeline execution.
  • Audit Logging - Maintains comprehensive logs of user actions and system changes to ensure compliance and visibility into operational history.
  • Data Workflow Integrations - Incorporates conversational models and lifecycle management tools directly into data workflows to automate complex processing.
  • Local Development Environments - Provides isolated interfaces for building, debugging, and mocking data workflows locally before production deployment.
  • Data Synchronization and Consistency - Ensures reliable and uniform data outputs across different sources and timing intervals to maintain a single source of truth.
  • User Identity Management - Integrates with enterprise identity providers to enforce role-based permissions and automate user provisioning through standard authentication protocols.
  • Legacy Migration Strategies - Converts existing workflow definitions into modern data assets while maintaining backward compatibility during system transitions.

Historique des stars

Graphique de l'historique des stars pour dagster-io/dagsterGraphique de l'historique des stars pour dagster-io/dagster

Recherche par IA

Explorez plus de dépôts awesome

Décrivez vos besoins en langage naturel — l'IA classe des milliers de projets open source sélectionnés par pertinence.

Start searching with AI

Questions fréquentes

Que fait dagster-io/dagster ?

Dagster is a data orchestration platform designed to manage the entire lifecycle of data assets through declarative modeling and version-controlled code. It functions as a workflow engine that treats data assets as first-class primitives, allowing teams to define, schedule, and monitor complex pipelines while maintaining clear visibility into lineage, dependencies, and data quality.

Quelles sont les fonctionnalités principales de dagster-io/dagster ?

Les fonctionnalités principales de dagster-io/dagster sont : Data Pipeline Orchestration, Workflow Orchestration Engines, Declarative Orchestration, Data Lineage, Data Asset Lifecycle Management, Data Asset Modeling, Data Pipeline Definitions, Configuration-as-Code Frameworks.

Quelles sont les alternatives open-source à dagster-io/dagster ?

Les alternatives open-source à dagster-io/dagster incluent : prefecthq/prefect — Prefect is a workflow orchestration platform designed to define, schedule, and monitor complex data pipelines as… datahub-project/datahub — DataHub is a metadata management platform designed to unify technical, operational, and business context across… apache/airflow — Airflow is a platform for programmatically authoring, scheduling, and monitoring complex data pipelines. It functions… apache/incubator-airflow — This project is a Python workflow orchestration platform and programmatic data pipeline engine used to author,… dbt-labs/dbt-core — dbt-core is a command-line framework for transforming data within a warehouse using modular SQL and version control.… spotify/luigi — Luigi is a Python framework designed for building and managing complex batch data pipelines. It functions as a…

Alternatives open source à Dagster

Projets open source similaires, classés selon le nombre de fonctionnalités partagées avec Dagster.
  • prefecthq/prefectAvatar de PrefectHQ

    PrefectHQ/prefect

    21,640Voir sur GitHub↗

    Prefect is a workflow orchestration platform designed to define, schedule, and monitor complex data pipelines as Python code. It functions as a container-native engine that wraps individual tasks in isolated environments, ensuring consistent dependencies and resource allocation across diverse infrastructure. By utilizing a state-machine-based orchestration model, the system tracks execution progress through discrete transitions and persistent event logs to maintain reliable and observable task processing. The platform distinguishes itself through a decoupled worker-API architecture, which sep

    Pythonautomationdatadata-engineering
    Voir sur GitHub↗21,640
  • datahub-project/datahubAvatar de datahub-project

    datahub-project/datahub

    12,141Voir sur GitHub↗

    DataHub is a metadata management platform designed to unify technical, operational, and business context across diverse data ecosystems. By utilizing a graph-based metadata model and an event-driven ingestion architecture, it creates a centralized source of truth that maps complex data relationships, lineage, and ownership. This foundational framework enables organizations to maintain a synchronized view of their data landscape, supporting both human-led discovery and automated data operations. The platform distinguishes itself through its focus on grounding artificial intelligence and autono

    Pythondata-catalogdata-discoverydata-governance
    Voir sur GitHub↗12,141
  • apache/airflowAvatar de apache

    apache/airflow

    45,902Voir sur GitHub↗

    Airflow is a platform for programmatically authoring, scheduling, and monitoring complex data pipelines. It functions as a workflow automation engine that manages the lifecycle of recurring business processes by executing code-defined task dependencies. By representing workflows as directed acyclic graphs, the system ensures that task execution order and data flow are explicitly defined and reliably maintained across distributed computing environments. The platform distinguishes itself through a highly modular, provider-based architecture that decouples core orchestration logic from external

    Pythonairflowapacheapache-airflow
    Voir sur GitHub↗45,902
  • apache/incubator-airflowAvatar de apache

    apache/incubator-airflow

    45,840Voir sur GitHub↗

    This project is a Python workflow orchestration platform and programmatic data pipeline engine used to author, schedule, and monitor complex data pipelines. It functions as a directed acyclic graph manager and scheduler, allowing users to define data movement and transformation tasks as code to ensure precise execution order and maintainability. The platform distinguishes itself by treating workflows as code, enabling pipelines to be versioned and tested through a standard programming language. It utilizes a system of extensible operators to encapsulate integration logic and employs a templat

    Python
    Voir sur GitHub↗45,840
Voir les 30 alternatives à Dagster→