12 Repos
Frameworks for executing multi-stage machine learning inference pipelines.
Distinguishing note: Focuses on pipeline orchestration rather than individual model execution.
Explore 12 awesome GitHub repositories matching artificial intelligence & ml · Inference Pipeline Orchestrators. Refine with filters or upvote what's useful.
Ray is a distributed computing framework designed to scale Python and Java applications across clusters by abstracting task scheduling and resource management. It functions as a resource-aware execution engine that manages task dependencies, placement, and fault tolerance across networked compute nodes. At its core, the system provides a stateful actor model, allowing developers to define classes that run in dedicated processes to maintain and mutate internal state across remote method calls. The framework distinguishes itself through a robust cross-language interoperability layer, enabling f
Executes multi-stage inference pipelines that handle preprocessing, tokenization, and accelerated GPU inference.
Chatterbox is a comprehensive machine learning platform designed for multilingual speech synthesis and real-time audio generation. It functions as an engine that converts text into natural-sounding speech, capable of replicating specific human vocal characteristics and emotional expressions from short audio samples. The platform distinguishes itself through advanced control over the synthesis process, allowing for the manipulation of emotional intensity and the injection of non-verbal vocalizations such as laughter or coughing. It is engineered for low-latency performance, utilizing an optimi
Orchestrates multi-stage inference pipelines to minimize latency for real-time voice applications.
WhisperX is an automated speech recognition toolkit designed to convert spoken audio into text while maintaining precise synchronization with the original media. It functions as an integrated pipeline that combines transcription, phoneme-based alignment, and speaker diarization to produce structured, attributed transcripts. The project distinguishes itself through its use of forced alignment, which matches existing text to audio signals at the phoneme level to generate accurate word-level timestamps. It also incorporates speaker diarization to identify and label unique voices within a recordi
Sequences multiple machine learning models into an integrated pipeline for transcription, alignment, and speaker identification.
Stable Diffusion WebUI Forge is a web-based interface and inference engine designed for the generation of AI media. It functions as a platform for executing diffusion-based models, providing a centralized environment to manage image preprocessors, custom generation logic, and hardware-accelerated sampling. The project distinguishes itself through a neural network patching framework that allows for the modification of model layers and the application of spatial conditioning during inference. By injecting custom logic and adapters directly into the network, users can influence output behaviors
Orchestrates model loading, memory allocation, and data processing sequences to ensure efficient execution on limited hardware.
Triton Inference Server is a high-performance AI model inference server and multi-framework model runtime designed for deploying machine learning models across cloud, data center, and embedded edge infrastructure. It serves as an execution engine that allows for the concurrent running of models from various frameworks to optimize hardware utilization. The project features a dynamic batching inference engine that groups individual requests into larger batches to increase total processing throughput. It also provides a model ensemble pipeline, which enables the chaining of multiple models toget
Provides a framework for executing multi-stage machine learning inference pipelines using model ensembles.
Triton Inference Server is a high-performance server designed to deploy machine learning models from multiple frameworks across GPUs and CPUs. It functions as a hardware-accelerated inference engine and a gRPC inference gateway, providing a standardized communication layer for transmitting binary tensor data with low latency. The system acts as a multi-framework model orchestrator, allowing users to link multiple AI models into ensembles and scripts to create complex inference pipelines. It also serves as a model lifecycle manager, providing controls to load, unload, and monitor the performan
Provides a system for linking multiple models into ensembles and scripts to create complex, multi-stage inference pipelines.
Apache Beam is a distributed data pipeline framework and unified data processing model designed to handle both bounded batch data and unbounded real-time streams. It provides a system for building scalable, data-parallel workflows that operate across compute clusters using a single programming model. The framework utilizes a cross-runner pipeline abstraction that decouples the data processing logic from the underlying execution backend, allowing the same pipeline to run on different distributed compute engines. It supports multi-language pipeline development by translating high-level code fro
Provides a framework for executing multi-stage machine learning inference pipelines during data transformations.
KServe is an open platform for deploying and serving generative and predictive AI models on Kubernetes. It defines inference services as custom resources with declarative YAML specifications, enabling a Kubernetes-native approach to model deployment and lifecycle management. The platform leverages Knative-based serverless scaling for automatic scale-to-zero and revision management, and supports a pluggable serving runtime architecture that maps model formats to containerized execution environments. KServe distinguishes itself through model-aware autoscaling that scales replicas based on token
Orchestrates complex inference workflows by chaining, ensembling, and routing through multiple models.
KServe is a Kubernetes-native platform for deploying and serving machine learning models as scalable inference services. It supports both generative AI models, including large language models, and traditional predictive models from frameworks such as TensorFlow, PyTorch, Scikit-Learn, XGBoost, and ONNX. The platform manages the full lifecycle of model deployments, including revision tracking, canary rollouts, A/B testing, and automatic rollbacks, and provides serverless scale-to-zero capabilities for cost-efficient resource management. KServe distinguishes itself through a standardized infere
Orchestrates complex inference workflows by chaining multiple models into ensembles, pipelines, and conditional routing graphs.
This project provides a comprehensive technical guide and framework for engineering large-scale machine learning systems. It covers the full lifecycle of model development, focusing on the infrastructure and computational principles required to build, train, and serve generative AI models across distributed GPU clusters. The repository distinguishes itself by offering deep-dive tutorials and implementation strategies for complex system challenges. It emphasizes high-performance architectural primitives, such as collective communication orchestration, distributed tensor sharding, and static gr
Provides a unified framework for chaining sequential inference tasks like audio encoding, reasoning, and synthesis.
Seldon Core ist ein auf Kubernetes basierender Server für Machine-Learning-Modelle und ein MLOps-Inference-Framework. Es fungiert als Serving-Engine für mehrere Modelle und als Pipeline-Orchestrator, der Modelle als skalierbare Microservices verpackt, die über standardisierte REST- und gRPC-APIs bereitgestellt werden. Das Projekt zeichnet sich durch graphbasierte Inference-Pipelines aus, die Modelle und Datentransformatoren zu sequenziellen Workflows verketten. Es optimiert die Hardwareauslastung durch Shared-Serving für mehrere Modelle und Strategien für dynamisches Memory-Overcommit, während es gleichzeitig Produktionsexperimente durch gewichtetes Traffic-Routing, A/B-Tests und Shadow-Deployments unterstützt. Das Framework deckt ein breites Spektrum an MLOps-Funktionen ab, darunter bedarfsgesteuertes Autoscaling, asynchrone Request-Verarbeitung über Message-Busse sowie umfassendes Monitoring für Data Drift, Ausreißer und die Erklärbarkeit von Vorhersagen. Es bietet zudem Infrastrukturmanagement für die Konfiguration der Modell-Runtime und sichere Kommunikation mittels TLS-Verschlüsselung über Control- und Data-Planes hinweg.
Orchestrates complex sequences of models and data transformers into multi-stage inference workflows.
Serving ist ein High-Performance-Framework für die Bereitstellung und Skalierung von Machine-Learning-Modellen als Produktionsdienste. Es fungiert als verteilte Inferenz-Engine, die die Ausführung komplexer Datenverarbeitungs-Workflows durch die Verkettung mehrerer Modelle in gerichteten azyklischen Graphen ermöglicht. Die Plattform zeichnet sich durch ihre Fähigkeit aus, den gesamten Lebenszyklus von Produktionsmodellen zu verwalten, was Hot-Swappable-Versioning ermöglicht, das Dienste ohne Ausfallzeiten aktualisiert. Sie unterstützt horizontale Skalierung durch verteiltes Modell-Sharding und optimiert den Abruf hochdimensionaler Daten durch spezialisierte Sparse-Parameter-Lookup-Strukturen. Das System bietet eine umfassende Suite an Funktionen für Produktionsumgebungen, einschließlich hardwarebeschleunigter Inferenz-Ausführung, mehrsprachiger RPC-Schnittstellen und integriertem Service-Monitoring. Es enthält zudem Sicherheitsfunktionen wie Request-Authentifizierung und verschlüsselte Kommunikationskanäle, um Modellbereitstellungen zu schützen.
Orchestrates multi-stage inference pipelines using directed graphs to manage data processing and prediction steps.