5 Repos
Creating, deleting, versioning, and inspecting pipeline definitions that transform incoming data before storage.
Distinct from Named Pipeline Ingestion: Distinct from Named Pipeline Ingestion: covers the full lifecycle of pipeline definitions (create, delete, version, inspect) rather than just referencing them during ingestion.
Explore 5 awesome GitHub repositories matching system administration & monitoring · Pipeline Lifecycle Managements. Refine with filters or upvote what's useful.
GreptimeDB is a distributed, open-source time-series database built for unified observability. It stores and queries metrics, logs, and traces together in a single columnar engine, supporting both SQL and PromQL for analysis. The database is designed as a Kubernetes-native operator with a decoupled compute and storage architecture, enabling horizontal scaling and multi-region deployment. What distinguishes GreptimeDB is its role as a multi-protocol ingestion gateway, accepting data through OpenTelemetry, Prometheus Remote Write, InfluxDB, Loki, Elasticsearch, Kafka, and MQTT protocols without
GreptimeDB removes a named pipeline and all its versions from the database through an HTTP interface.
ZenML is an extensible machine learning orchestration framework designed to manage the end-to-end lifecycle of data pipelines and AI agent workflows. It functions as a durable orchestrator that executes machine learning tasks as directed acyclic graphs, ensuring that every step is containerized for consistent performance across local, cloud, and hybrid infrastructure. By decoupling pipeline code from underlying compute and storage backends, the platform allows developers to define infrastructure-agnostic stacks that remain portable across diverse environments. The project distinguishes itself
Manages the full lifecycle of pipeline definitions, including deletion of obsolete runs to maintain an organized development environment.
ZenML is an orchestration platform designed for building, deploying, and monitoring reproducible machine learning pipelines and agentic workflows. It provides a unified framework that manages the entire lifecycle of machine learning assets, from data processing and model training to the deployment of persistent inference services. By decoupling pipeline logic from underlying compute and storage, the platform enables teams to transition workflows seamlessly from local development environments to production-grade cloud infrastructure. The platform distinguishes itself through a service-oriented
Manages the full lifecycle of pipeline definitions, including creation, deletion, and versioning.
Dieses Projekt ist ein Software Development Kit (SDK) und Cluster-Management-Tool für PHP. Es dient als SDK für Volltextsuche und Vektor-Suchschnittstelle, wodurch Anwendungen lexikalische, Fuzzy- und semantische Suchen auf indizierten Daten durchführen können. Die Bibliothek implementiert einen PSR-7-HTTP-Client, um die Kompatibilität zwischen verschiedenen Umgebungen durch standardisierte Messaging-Schnittstellen zu gewährleisten. Sie bietet eine spezialisierte Schnittstelle zum Abrufen von Embeddings und zur Durchführung semantischer Retrieval-Workflows unter Verwendung von Vektordaten. Der Funktionsumfang deckt eine breite Palette administrativer und operativer Aufgaben ab, einschließlich der Verwaltung von Suchindizes, der Überwachung des Cluster-Status und der Verwaltung von Dokumentlebenszyklen. Es unterstützt diverse Abfragemethoden wie SQL, EQL und ES|QL sowie Datenaggregation und Geodatenanalyse. Zusätzlich bietet es Tools für Machine-Learning-Orchestrierung, Anomalieerkennung sowie Identitäts- und Zugriffsmanagement.
Facilitates the creation and configuration of data pipelines for central log processing management.
Dag-factory ist ein Framework zur Erstellung und Verwaltung von Apache Airflow-Datenpipelines durch deklarative Konfigurationsdateien. Durch den Ersatz von manuellem prozeduralem Code durch strukturierte YAML-Definitionen ermöglicht es die programmatische Generierung komplexer Workflow-Strukturen, Task-Abhängigkeiten und Ausführungspläne. Das Projekt zeichnet sich dadurch aus, dass Konfigurationsschlüssel direkt auf Python-Klassenkonstruktoren und Operatoren abgebildet werden, was die dynamische Instanziierung von Objekten und benutzerdefinierter Logik ermöglicht. Es unterstützt hierarchische Konfigurationsvererbung zur Standardisierung von Einstellungen über Umgebungen hinweg und bietet Mechanismen zur direkten Injektion von Kubernetes-Pod-Spezifikationen in Task-Definitionen, um eine isolierte, skalierbare Ausführung zu gewährleisten. Das Framework deckt den gesamten Pipeline-Lebenszyklus ab, einschließlich automatisierter Dateierkennung, dynamischem Mapping auf Task-Ebene für parallele Verarbeitung und das Anhängen von Metadaten für die Integration externer Systeme. Es enthält zudem CLI-Tools zur Validierung von Konfigurationen, zum Auslösen von Ausführungen und zur Verwaltung von Umgebungsmigrationen.
Manages the full lifecycle of data pipelines, including validation, organization, and deployment across environments.