2 Repos
Centralized interfaces for tracking the state, artifacts, and execution of machine learning training jobs and pipelines.
Distinguishing note: None of the candidates cover a unified MLOps management interface; they focus on UI component lifecycles or binary interfaces.
Explore 2 awesome GitHub repositories matching artificial intelligence & ml · MLOps Control Planes. Refine with filters or upvote what's useful.
ClearML is a comprehensive MLOps platform designed to manage the entire machine learning lifecycle. It functions as an experiment tracking tool, a data versioning system, and a pipeline orchestrator, while providing infrastructure for GPU cluster management and model serving. The platform is distinguished by its ability to handle hybrid-cloud compute scheduling and fractional GPU allocation, allowing multiple workloads to share a single hardware accelerator. It employs a metadata-based approach to data versioning, using virtual views to track large datasets and artifacts without duplicating r
Provides a unified interface for tracking training jobs, pipelines, and model artifacts throughout their lifecycle.
TransformerLab ist eine MLOps-Orchestrierungsplattform und Forschungsumgebung, die für das Training, Fine-Tuning und die Evaluierung von Large Language Models entwickelt wurde. Sie dient als zentralisierte Steuerungsebene für das Management von Machine-Learning-Jobs und die Koordination verteilter GPU-Rechenleistung über hybride Cloud- und On-Premise-Anbieter hinweg. Die Plattform zeichnet sich durch agentengesteuerte Modelloptimierung aus und nutzt KI-Assistenten, um Metriken zu analysieren und automatisch Hyperparameter-Experimente vorzuschlagen und in die Warteschlange einzureihen. Sie bietet eine Remote-Entwicklungsumgebung, die es Benutzern ermöglicht, interaktive Notebooks, Code-Editoren und Secure-Shell-Sitzungen direkt auf Remote-Rechenknoten zu starten. Das System deckt ein breites Spektrum an Machine-Learning-Workflow-Funktionen ab, einschließlich verteilter Aufgabenkoordination, automatisierter Hyperparameter-Sweeps und umfassendem Experiment-Tracking. Es verfügt über integrierte Registries für die Versionierung von Datensätzen und Modell-Artefakten sowie Tools für die Evaluierung der Modell-Performance und das Deployment von Inference-Servern. Ein Command-Line-Interface wird für die Plattformsteuerung, das Job-Monitoring sowie die Verwaltung der Installation und Updates der lokalen Serverinstanz bereitgestellt.
Ships a centralized control plane for submitting and monitoring machine learning jobs across diverse compute providers.