18 dépôts
Calculates task execution order by treating workflows as directed acyclic graphs.
Distinct from Dependency Resolution: None of the candidates cover general DAG-based task orchestration; they focus on module imports or specific UI/platform dependencies.
Explore 18 awesome GitHub repositories matching software engineering & architecture · DAG-Based Dependency Resolution. Refine with filters or upvote what's useful.
Airflow is a workflow orchestration platform for authoring, scheduling, and monitoring complex data pipelines as code using Python. It employs a DAG-based task scheduler to manage execution timing and dependencies via directed acyclic graphs, utilizing a distributed task execution engine to run workloads across a cluster of worker nodes. The platform provides a data pipeline monitor for tracking the health and execution history of programmatic workflows. This includes a web interface for workflow progress visualization and health monitoring to identify and troubleshoot pipeline failures. The
Uses Directed Acyclic Graphs to determine the exact execution order and dependency mapping of complex workflows.
This project is a Python workflow orchestration platform and programmatic data pipeline engine used to author, schedule, and monitor complex data pipelines. It functions as a directed acyclic graph manager and scheduler, allowing users to define data movement and transformation tasks as code to ensure precise execution order and maintainability. The platform distinguishes itself by treating workflows as code, enabling pipelines to be versioned and tested through a standard programming language. It utilizes a system of extensible operators to encapsulate integration logic and employs a templat
Implements a DAG engine to determine the precise execution order of interdependent pipeline tasks.
DVC is a data versioning tool and pipeline orchestrator designed to track large datasets and machine learning models. It functions as a system for managing large data artifacts by storing lightweight metadata in version control while keeping the actual binaries in a separate cache. The project serves as an experiment tracker and remote storage synchronizer, enabling the execution and comparison of machine learning iterations based on hyperparameters and performance metrics. It provides a bridge for pushing and pulling these large data artifacts between local environments and cloud or on-premi
Provides DAG-based pipeline execution to orchestrate data processing steps and optimize re-execution.
Kedro is a data science pipeline framework and orchestration tool designed to build reproducible and modular data engineering workflows. It functions as an MLOps project template and Python data workflow tool that enforces software engineering best practices to move projects from prototype to production. The system distinguishes itself through a centralized data catalog manager that abstracts data access and versioning across various file formats and cloud storage systems. It further separates processing logic from data access via a lazy-loading data registry and provides a standardized proje
Determines task execution order by mapping function inputs and outputs to a directed acyclic graph.
Metaflow is a Python machine learning framework and MLOps workflow orchestrator designed to manage the lifecycle of data pipelines from local prototyping to production. It serves as a distributed compute manager and an experiment tracking system, enabling the creation of reproducible pipelines that transition between development and high-availability production environments. The framework distinguishes itself through an integrated checkpointing system that automatically persists intermediate data artifacts to remote storage, allowing failed runs to be resumed from the last successful step. It
Structures pipeline execution as a directed acyclic graph of steps with support for conditional branching and parallel execution.
Enterprise job scheduling middleware with distributed computing ability.
Models complex job pipelines as directed acyclic graphs to enforce execution order and data dependencies.
Hatchet is an open-source durable workflow engine and task orchestration platform. It provides a framework for building and executing fault-tolerant, multi-step pipelines as directed acyclic graphs (DAGs), with automatic retries, scheduling, and real-time observability. The system is built around durable task checkpointing, which persists execution state after each step so work can resume from the last checkpoint after a worker crash or restart, and it supports event-driven task resumption that pauses a task until a matching external event arrives. The platform distinguishes itself through it
Executes workflows as directed acyclic graphs with automatic parallelism and state persistence.
Open Multi-Agent is a TypeScript framework for multi-agent orchestration that decomposes natural language goals into a runtime-generated directed acyclic graph of tasks. It functions as a task orchestrator and workflow state manager, coordinating multiple AI models to execute parallel and sequential operations. The framework is distinguished by a proposer-judge consensus protocol used to validate agent outputs through a quorum of agreement. It employs provider-agnostic model routing to assign specific models to tasks based on roles or execution phases and utilizes state-based workflow checkpo
Decomposes natural language goals into directed acyclic graphs for parallel and sequential task execution.
Osmedeus is a security workflow orchestration engine that coordinates AI agents, shell commands, and scanning tools through declarative YAML pipelines. It functions as a distributed security scanner, a declarative workflow automator, and an AI agent framework for security, enabling automated multi-step security analysis with conditional branching, parallel execution, and distributed workers. The engine distinguishes itself through a hybrid runner model that executes workflow steps on the local host, inside Docker containers, or over SSH to remote machines, selected per step or module. It supp
Displays a graphical representation of workflow steps and their connections using an interactive flow editor.
Volcano is a Kubernetes-native batch scheduler specialized for AI, machine learning, and high-performance computing workloads. It provides gang scheduling to atomically allocate resources for all tasks of a distributed job, preventing deadlocks from partial allocation, and supports hierarchical queue management for multi-tenant resource isolation with configurable quotas, borrowing, and preemption. Topology-aware placement optimizes communication-intensive workloads by modeling network hierarchy to minimize cross-switch latency. Volcano differentiates itself with automated orchestration of di
Volcano defines lightweight Directed Acyclic Graph workflows for batch jobs with monitoring and validation.
Azkaban est un gestionnaire de flux de travail distribué et un orchestrateur de tâches basé sur DAG, conçu comme un processeur de batch d'entreprise. Il sert de moteur de flux de travail basé sur Java qui planifie et exécute des séquences de tâches complexes à travers un cluster de serveurs d'exécution, avec une fonctionnalité spécifique pour gérer les charges de travail big data sur les clusters Hadoop. Le système se distingue par un modèle d'exécuteur distribué qui coordonne l'état via une base de données partagée pour garantir une haute disponibilité. Il utilise une architecture basée sur des plugins qui permet des types de tâches personnalisés et des extensions de fonctionnalité système, y compris la possibilité de recharger à chaud les plugins sans redémarrer les serveurs d'exécution. La plateforme couvre un large éventail de capacités, y compris l'orchestration de pipelines de données avec logique conditionnelle, la planification périodique et basée sur des événements, et la surveillance d'entreprise avec suivi des SLA. Elle fournit un contrôle d'accès granulaire et l'usurpation d'identité utilisateur pour une exécution sécurisée, ainsi que des outils de gestion du trafic pour l'équilibrage de charge des exécuteurs et les quotas de ressources. Les utilisateurs peuvent gérer les flux de travail via une interface web ou par programmation via une API d'exécution de flux de travail.
Provides a graphical representation of the workflow showing the relationships and dependencies between jobs.
ms-agent is an LLM agent framework and multi-agent orchestration system designed to build autonomous entities that combine large language models with tool calling and structured workflows. It serves as a tool integration platform and workflow engine for executing complex tasks through the coordination of specialized agents. The project distinguishes itself through a multimodal agent workflow engine capable of automating the production of text, images, and video. It features a sandboxed code execution environment for running generated code and quantitative data analysis in isolated containers,
Uses directed acyclic graphs to map dependencies between agent skills and ensure correct execution order.
OpenSquilla est un framework d'orchestration d'agents LLM conçu pour coordonner des workflows IA multi-étapes et l'exécution d'outils via des graphes orientés acycliques (DAG). Il fonctionne comme un système centralisé pour gérer des packages de compétences spécialisés et exécuter des séquences de raisonnement complexes. Le projet se distingue par une passerelle de routage qui dirige les tâches vers différents fournisseurs d'IA en fonction de la complexité, du coût et de la performance. Il utilise un système de mémoire IA à plusieurs niveaux qui organise les connaissances de travail, épisodiques et sémantiques à l'aide d'embeddings locaux et de SQLite, ainsi qu'un bac à sable d'exécution sécurisé qui isole le code généré par l'agent via des profils de permission basés sur les risques. La plateforme couvre un large éventail de capacités, incluant le déploiement multicanal vers le web et les plateformes de messagerie, la planification automatisée des tâches via cron, et un pont Model Context Protocol pour se connecter à des outils externes. Elle fournit également des outils complets de surveillance et d'observabilité pour suivre les coûts en jetons, auditer les décisions d'exécution et gérer un catalogue de compétences réutilisables. Le système inclut des utilitaires en ligne de commande pour l'initialisation de l'espace de travail et la gestion du cycle de vie des compétences.
Coordinates complex reasoning steps and tool dependencies using directed acyclic graphs to manage multi-step AI workflows.
Ce projet est un framework d'automatisation de navigateur LLM et une interface de navigateur pour agent IA. Il sert de couche de contrôle qui traduit les instructions en langage naturel en interactions de navigateur en utilisant des grands modèles de langage, permettant aux agents IA de naviguer et d'interagir avec des pages web via des fonctions de contrôle de navigateur standardisées. Le système fonctionne comme un orchestrateur de flux de travail RPA et un outil de gestion de navigateur headless, capable d'enregistrer et de rejouer des séquences de navigateur déterministes pour automatiser des tâches répétitives. Il se distingue par des configurations furtives, incluant des proxies résidentiels et des moteurs de navigateur modifiés, pour contourner la détection de bots et résoudre les CAPTCHAs. La plateforme couvre un large éventail de capacités, notamment l'extraction structurée de données web, la gestion de session persistante pour maintenir l'authentification et l'intervention humaine pour les étapes complexes comme l'authentification multi-facteurs. Il prend en charge à la fois la connectivité locale et les déploiements en sandbox cloud gérés, offrant une gestion visuelle des flux de travail et une surveillance de l'activité en temps réel via des graphiques interactifs. L'intégration est fournie via une interface en ligne de commande et une connectivité API pour les fournisseurs de LLM externes et les plateformes d'orchestration tierces.
Offers a graphical interface to visualize automation workflows as interactive graphs with real-time execution logs.
EFCore.BulkExtensions is a library for executing high-performance batch insert, update, and delete operations within the Entity Framework Core ecosystem. It functions as a database batch processing toolkit and a wrapper for native SQL Bulk Copy to enable faster data ingestion and synchronization across multiple database providers. The library provides specialized capabilities for relational data synchronization, allowing users to align database tables with local entity lists through bulk upserts and conditional synchronization. It also supports relational data graph insertions, which enable t
Analyzes entity relationships using directed acyclic graphs to determine the correct order for bulk operations.
Polyaxon is a Kubernetes-native machine learning orchestration platform and MLOps pipeline orchestrator. It serves as a control plane for managing distributed deep learning workloads, automated machine learning pipelines, and experiment tracking. The platform distinguishes itself through specialized services for distributed training management, including MPI-based coordination for PyTorch and TensorFlow. It provides an automated hyperparameter optimization service utilizing Bayesian, random, and grid search algorithms, alongside managed interactive AI workspaces for launching Jupyter notebook
Coordinates interdependent training and evaluation tasks by executing them as a directed acyclic graph.
Nuke is a build automation system for defining software compilation and deployment pipelines using a strongly typed C# console application. It functions as a cross-platform build engine and pipeline orchestrator that treats build configurations as standard executable programs rather than static files. By leveraging a compiled language, the system provides type safety and IDE support for build script logic. This approach allows for the definition of automation and CI/CD pipelines using a professional programming language instead of YAML or shell scripts. The engine manages .NET project orches
Calculates the correct build sequence by treating targets as nodes in a directed acyclic graph.
Dag-factory est un framework pour construire et gérer des pipelines de données Apache Airflow via des fichiers de configuration déclaratifs. En remplaçant le code procédural manuel par des définitions YAML structurées, il permet la génération programmatique de structures de workflow complexes, de dépendances de tâches et de calendriers d'exécution. Le projet se distingue en mappant les clés de configuration directement aux constructeurs de classes et opérateurs Python, permettant l'instanciation dynamique d'objets et une logique personnalisée. Il prend en charge l'héritage de configuration hiérarchique pour standardiser les paramètres entre les environnements et fournit des mécanismes pour injecter des spécifications de pods Kubernetes directement dans les définitions de tâches afin d'assurer une exécution isolée et évolutive. Le framework couvre l'intégralité du cycle de vie du pipeline, incluant la découverte automatique de fichiers, le mappage dynamique au niveau des tâches pour le traitement parallèle et l'attachement de métadonnées pour l'intégration avec des systèmes externes. Il inclut également des utilitaires en ligne de commande pour valider les configurations, déclencher des exécutions et gérer les migrations d'environnement.
Constructs and executes directed acyclic graphs of tasks programmatically at runtime based on configuration definitions.