9 रिपॉजिटरी
Tools for versioning, quality screening, and managing data pipelines.
Explore 9 awesome GitHub repositories matching part of an awesome list · Data Orchestration. Refine with filters or upvote what's useful.
Pathway is a high-performance data processing framework designed for building unified batch and streaming pipelines. It functions as an orchestrator for complex data transformations, utilizing a differential dataflow engine to process updates incrementally. By treating static datasets and continuous event streams with identical logic, the platform ensures exactly-once processing semantics and consistent results across diverse data sources. The framework distinguishes itself through its specialized support for real-time artificial intelligence and retrieval-augmented generation. It features in
Performant Python ETL framework with a Rust runtime.
Cadence is a distributed workflow orchestration engine designed to execute long-running, asynchronous business logic with built-in durability and resilience across distributed systems. It functions as a stateful process manager that ensures processes resume from their last known state following system crashes or network outages. The platform utilizes a distributed task queue to manage work across independent worker nodes and supports persistence via SQL or Cassandra backend storage. It includes a workflow visualization dashboard for inspecting execution histories and state traces, alongside a
Scalable orchestration engine for durable distributed workflows.
lakeFS is a data lake versioning system that provides Git-like branching and commits for large datasets stored in object storage. It functions as a version control layer, enabling the creation of immutable snapshots, atomic commits, and zero-copy branching to create isolated environments for data experimentation without duplicating physical files. The system serves as an S3-compatible storage gateway and an Iceberg REST catalog, allowing standard cloud storage protocols and compatible clients to manage versioned tables. It acts as a data quality gatekeeper by using an event-driven hook system
Version control for data lakes on object storage.
Build data pipelines with SQL and Python, ingest data from different sources, add quality checks, and build end-to-end flows.
Data pipeline framework supporting SQL and Python DAGs.
Git-like version control for data tables and views.
Substation is a toolkit for routing, normalizing, and enriching security event and audit logs.
Cloud-native data pipeline and transformation toolkit.
Model Context Protocol (MCP) Server for the Keboola Platform
Facilitates data workflow management and analytics integration.
Real-time data quality screening API for pipeline ingestion.