awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
dagster-io avatar

dagster-io/dagster

0
View on GitHub↗
14,974 stars·1,986 forks·Python·apache-2.0·83 viewsdagster.io↗

Dagster

Dagster is a data orchestration platform designed to manage the entire lifecycle of data assets through declarative modeling and version-controlled code. It functions as a workflow engine that treats data assets as first-class primitives, allowing teams to define, schedule, and monitor complex pipelines while maintaining clear visibility into lineage, dependencies, and data quality.

The platform distinguishes itself by using a code-as-configuration framework that enables standard software engineering practices, such as unit testing and local mocking, to be applied directly to data workflows. Its architecture is built on a pluggable execution engine that decouples orchestration logic from the underlying compute, allowing tasks to run across diverse cloud-native, serverless, and containerized environments. Furthermore, it supports partition-aware scheduling, which enables incremental processing and efficient management of high-volume datasets.

Beyond core orchestration, the system provides a comprehensive suite of tools for data platform management, including automated quality governance, infrastructure cost optimization, and centralized asset cataloging. It integrates with enterprise identity providers for access control and offers robust observability features, such as streaming logs and visual lineage tracking, to ensure system health and compliance.

The platform supports a variety of deployment models, ranging from self-hosted and hybrid configurations to a fully managed control plane. It includes specialized utilities for migrating legacy pipelines and operationalizing interactive scripts into production-ready components.

Features

  • Data Pipeline Orchestration - Builds and schedules complex data workflows as version-controlled code to ensure reliable execution.
  • Workflow Orchestration Engines - Coordinates distributed tasks and data dependencies across heterogeneous cloud environments and external infrastructure.
  • Declarative Orchestration - Models data pipelines as a graph of versioned assets where the system automatically determines execution order based on dependency requirements.
  • Data Lineage - Provides visual and searchable mapping of data flows to document relationships between source inputs and final downstream assets.
  • Data Asset Lifecycle Management - Models data assets as first-class primitives to manage lineage, dependencies, and quality from ingestion through delivery.
  • Data Asset Modeling - Models data as declarative assets to track lineage, quality, and dependencies throughout the entire lifecycle.
  • Data Pipeline Definitions - Expresses data workflows, resources, and scheduling logic as version-controlled code to enable standard software engineering practices.
  • Configuration-as-Code Frameworks - Expresses entire data workflows and infrastructure resources as version-controlled code to enable standard software engineering practices.
  • Observability Platforms - Monitors the health, performance, and lineage of data workflows and assets through a unified interface.
  • Data Ingestion - Provides modular tools to ingest and transform raw information from various sources into usable assets for downstream analysis.
  • Data Partitioning - Divides large datasets into logical slices to enable incremental processing, targeted re-runs, and efficient management of high-volume workflows.
  • Data Quality Frameworks - Integrates validation checks and automated policies directly into pipelines to ensure data integrity throughout the asset lifecycle.
  • Data Catalogs - Maintains a centralized view of data assets, workflows, and lineage to help teams discover and reuse components.
  • Orchestration Engines - Decouples the orchestration logic from the compute layer to allow tasks to run across diverse environments like containers or serverless.
  • Managed Control Planes - Provides a hosted control plane with enterprise-grade features and insights for teams requiring fully managed infrastructure.
  • Infrastructure as Code - Defines deployment environments and infrastructure configurations using declarative files to integrate with existing workflows.
  • Infrastructure as Code Practices - Manages data pipelines and infrastructure configurations using software engineering practices like testing and CI/CD.
  • Data Access Governance - Centralizes visibility across teams and environments while enforcing fine-grained access controls and maintaining audit-ready lineage for compliance.
  • Automated Quality Workflows - Embeds validation checks and automated testing directly into data pipelines to ensure data integrity.
  • Pipeline Observability Tools - Streams execution logs and performance metrics to external monitoring platforms to maintain unified observability and system health.
  • Data Analysis and Processing - Orchestration for data assets and pipelines.
  • Data Pipelines - Orchestration platform for data development and production execution.
  • Data Pipelines and Orchestration - Orchestration platform for developing and observing data assets.
  • Data Processing - Data orchestrator for ML and ETL workflows.
  • Workflow Orchestration - Library for building and orchestrating data-intensive applications.
  • Data Engineering - Data orchestrator for machine learning and ETL workflows.
  • Data Orchestration - Data orchestrator for managing complex machine learning workflows.
  • Job Schedulers - Orchestration platform for data assets.
  • Scheduling - Orchestrates data pipelines for analytics and machine learning.
  • Workflow Orchestration - Development, production and observation of data assets.
  • Workflow Scheduling - Data orchestrator for machine learning and ETL pipelines.
  • General Purpose Orchestration - Data orchestrator for machine learning, analytics, and ETL pipelines.
  • Workflow Frameworks - API for defining DAGs to build data-intensive applications.
  • Workflow Scheduling - Data orchestrator designed for machine learning, analytics, and ETL pipelines.
  • Data Storage - Connects to cloud object stores, data warehouses, and lakehouse architectures to read, write, and version data assets securely.
  • Event-Driven Data Pipelines - Executes data tasks based on custom schedules or external events to automate pipeline runs in response to real-time data changes.
  • Distributed Computing - Launches and manages code execution across cloud-native, serverless, and containerized infrastructure to scale processing power.
  • Deployment Environments - Supports self-hosted, hybrid, and managed infrastructure deployment models to avoid vendor lock-in.
  • Resource Cost Management - Tracks and attributes resource consumption and compute expenses to specific data assets and pipeline runs.
  • Secret Management - Retrieves and rotates sensitive API keys and configuration parameters from secure vaults automatically during pipeline execution.
  • Execution Metadata - Captures and indexes execution logs, lineage, and asset state in a centralized store to provide unified visibility.
  • Automated Alerting Workflows - Notifies teams immediately when data quality checks fail or budget thresholds are exceeded.
  • Unit Testing - Verifies the logic and reliability of data processing code using unit, integration, and mock testing frameworks before deployment.
  • Isolated Execution Environments - Creates ephemeral, production-mirroring sandboxes for pull requests to validate changes end-to-end.
  • Local Development Tools - Provides isolated interfaces and testing frameworks that allow developers to simulate production environments and validate pipeline logic locally.
  • Production Operationalization - Converts interactive notebooks and shell scripts into production-ready pipeline components to eliminate manual deployment overhead.
  • Cloud Storage - Reads and writes data objects to cloud storage buckets to provide a resilient persistence layer for pipeline execution.
  • Audit Logging - Maintains comprehensive logs of user actions and system changes to ensure compliance and visibility into operational history.
  • Data Workflow Integrations - Incorporates conversational models and lifecycle management tools directly into data workflows to automate complex processing.
  • Local Development Environments - Provides isolated interfaces for building, debugging, and mocking data workflows locally before production deployment.
  • Data Synchronization and Consistency - Ensures reliable and uniform data outputs across different sources and timing intervals to maintain a single source of truth.
  • User Identity Management - Integrates with enterprise identity providers to enforce role-based permissions and automate user provisioning through standard authentication protocols.
  • Legacy Migration Strategies - Converts existing workflow definitions into modern data assets while maintaining backward compatibility during system transitions.

Star history

Star history chart for dagster-io/dagsterStar history chart for dagster-io/dagster

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with Dagster

These projects share indexed features with Dagster. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • prefecthq/prefectPrefectHQ avatar

    PrefectHQ/prefect

    21,640View on GitHub↗

    Prefect is a workflow orchestration platform designed to define, schedule, and monitor complex data pipelines as Python code. It functions as a container-native engine that wraps individual tasks in isolated environments, ensuring consistent dependencies and resource allocation across diverse infrastructure. By utilizing a state-machine-based orchestration model, the system tracks execution progress through discrete transitions and persistent event logs to maintain reliable and observable task processing. The platform distinguishes itself through a decoupled worker-API architecture, which sep

    Pythonautomationdatadata-engineering
    View on GitHub↗21,640
  • datahub-project/datahubdatahub-project avatar

    datahub-project/datahub

    12,141View on GitHub↗

    DataHub is a metadata management platform designed to unify technical, operational, and business context across diverse data ecosystems. By utilizing a graph-based metadata model and an event-driven ingestion architecture, it creates a centralized source of truth that maps complex data relationships, lineage, and ownership. This foundational framework enables organizations to maintain a synchronized view of their data landscape, supporting both human-led discovery and automated data operations. The platform distinguishes itself through its focus on grounding artificial intelligence and autono

    Pythondata-catalogdata-discoverydata-governance
    View on GitHub↗12,141
  • apache/airflowapache avatar

    apache/airflow

    45,902View on GitHub↗

    Airflow is a platform for programmatically authoring, scheduling, and monitoring complex data pipelines. It functions as a workflow automation engine that manages the lifecycle of recurring business processes by executing code-defined task dependencies. By representing workflows as directed acyclic graphs, the system ensures that task execution order and data flow are explicitly defined and reliably maintained across distributed computing environments. The platform distinguishes itself through a highly modular, provider-based architecture that decouples core orchestration logic from external

    Pythonairflowapacheapache-airflow
    View on GitHub↗45,902
  • apache/incubator-airflowapache avatar

    apache/incubator-airflow

    45,840View on GitHub↗

    This project is a Python workflow orchestration platform and programmatic data pipeline engine used to author, schedule, and monitor complex data pipelines. It functions as a directed acyclic graph manager and scheduler, allowing users to define data movement and transformation tasks as code to ensure precise execution order and maintainability. The platform distinguishes itself by treating workflows as code, enabling pipelines to be versioned and tested through a standard programming language. It utilizes a system of extensible operators to encapsulate integration logic and employs a templat

    Python
    View on GitHub↗45,840
Compare all 30 related projects→

Frequently asked questions

What does dagster-io/dagster do?

Dagster is a data orchestration platform designed to manage the entire lifecycle of data assets through declarative modeling and version-controlled code. It functions as a workflow engine that treats data assets as first-class primitives, allowing teams to define, schedule, and monitor complex pipelines while maintaining clear visibility into lineage, dependencies, and data quality.

What are the main features of dagster-io/dagster?

The main features of dagster-io/dagster are: Data Pipeline Orchestration, Workflow Orchestration Engines, Declarative Orchestration, Data Lineage, Data Asset Lifecycle Management, Data Asset Modeling, Data Pipeline Definitions, Configuration-as-Code Frameworks.

Which projects share features with dagster-io/dagster?

Projects with overlapping indexed features include: prefecthq/prefect — Prefect is a workflow orchestration platform designed to define, schedule, and monitor complex data pipelines as… datahub-project/datahub — DataHub is a metadata management platform designed to unify technical, operational, and business context across… apache/airflow — Airflow is a platform for programmatically authoring, scheduling, and monitoring complex data pipelines. It functions… apache/incubator-airflow — This project is a Python workflow orchestration platform and programmatic data pipeline engine used to author,… dbt-labs/dbt-core — dbt-core is a command-line framework for transforming data within a warehouse using modular SQL and version control.… spotify/luigi — Luigi is a Python framework designed for building and managing complex batch data pipelines. It functions as a…