awesome-repositories.com
Blog
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoAcerca deCómo clasificamosPrensaServidor MCP
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
polyaxon avatar

polyaxon/polyaxon

0
View on GitHub↗
3,707 estrellas·325 forks·MDX·Apache-2.0·12 vistaspolyaxon.com↗

Polyaxon

Polyaxon is a Kubernetes-native machine learning orchestration platform and MLOps pipeline orchestrator. It serves as a control plane for managing distributed deep learning workloads, automated machine learning pipelines, and experiment tracking.

The platform distinguishes itself through specialized services for distributed training management, including MPI-based coordination for PyTorch and TensorFlow. It provides an automated hyperparameter optimization service utilizing Bayesian, random, and grid search algorithms, alongside managed interactive AI workspaces for launching Jupyter notebooks and visualization tools.

Its broader capabilities cover the end-to-end model lifecycle, including container-native directed acyclic graph workflow automation, artifact and data lineage tracking, and a centralized model registry. The system also includes tools for compute resource management across multi-cluster environments and role-based access control for team collaboration.

The platform is deployed and managed using Helm charts for automated configuration on Kubernetes.

Features

  • Machine Learning Orchestration - Serves as a control plane for managing and orchestrating large-scale distributed machine learning workloads on Kubernetes.
  • Distributed Training - Coordinates large-scale deep learning workloads across multiple compute nodes using a message passing interface.
  • Centralized Control Planes - Provides a centralized control plane to coordinate the lifecycle, scaling, and configuration of distributed ML workloads.
  • Kubernetes Orchestrators - Orchestrates deep learning and GPU-accelerated workloads using native Kubernetes resources and operators.
  • Distributed ML Workload Deployment - Orchestrates deep learning jobs across single-node or multi-node Kubernetes environments to ensure scalability.
  • AI Workload Orchestration - Provides a programmable interface to manage the submission and execution of deep learning and GPU-accelerated workloads.
  • Data Lineage - Maintains a provenance record of typed inputs and outputs to trace the history of data throughout the ML workflow.
  • Distributed Training Managers - Coordinates multi-node deep learning experiments using MPI and other distributed compute patterns.
  • Experiment Tracking - Records metrics, artifacts, and code versions to ensure reproducibility and analyze model performance across experiments.
  • Experiment Tracking Systems - Implements a registry for versioning models and logging metrics, hyperparameters, and artifacts to ensure reproducibility.
  • Experiment Visualization Dashboards - Provides dedicated interfaces for comparing and analyzing logged machine learning experiment metrics and results via a visualization service.
  • Experiment Metadata Tracking - Provides a system for recording hyperparameters, performance metrics, and version history to ensure scientific reproducibility of AI experiments.
  • Interactive AI Workspace Providers - Provides a managed environment for launching Jupyter notebooks and visualization tools integrated with compute clusters.
  • Interactive AI Workspaces - Provisions managed Jupyter notebooks and visualization tools integrated with compute clusters for collaborative research.
  • Distributed TensorFlow Orchestration - Coordinates distributed training across multiple nodes using a structured replica interface for TensorFlow workers and servers.
  • Model Lifecycle Management - Manages the end-to-end model lifecycle, including versioning in a central registry and tracking data lineage.
  • Containerized Training - Executes data processing and training tasks in containerized environments to ensure consistent and reproducible runtimes.
  • ML Workflow Automation - Automates the iterative cycle of data synthesis, model training, and evaluation using DAGs.
  • ML Workflow Orchestrations - Coordinates the end-to-end process of researching, coding, training, and deploying models via distributed scheduling.
  • Model Registries - Provides a centralized registry for versioning trained models and managing their lifecycle for deployment.
  • Hyperparameter Optimization - Runs iterative processes that generate and tune hyperparameters to optimize overall model performance.
  • Hyperparameter Tuning - Automates the search for optimal model parameters using Bayesian Optimization or Grid Search.
  • Distributed PyTorch Orchestration - Coordinates distributed training experiments for PyTorch by managing master and worker replicas across a compute cluster.
  • ML Pipeline Automation - Automates sequences of interdependent machine learning tasks using container-native directed acyclic graphs.
  • Artifact Versioning - Tracks the provenance and versions of large binary models and datasets to ensure reproducibility.
  • Interactive Notebook Environments - Launches managed Jupyter notebooks within project environments for interactive development and exploration.
  • Kubernetes Workspace Provisioning - Provisions managed development environments as Kubernetes pods, including interactive notebooks and editors.
  • Iterative Parameter Optimization - Runs automated search algorithms like Bayesian or Random search to iteratively optimize model performance metrics.
  • Dataset Versioning - Tracks and locks versions of machine learning datasets and the specific execution runs that generated them.
  • MLOps Pipeline Automation - Defines and executes container-native directed acyclic graphs to automate professional machine learning workflows.
  • Cloud Infrastructure Scaling - Deploys jobs and experiments across multiple cloud providers or on-premises hardware to manage concurrency.
  • ML Workload Schedulers - Schedules machine learning training and evaluation tasks across compute clusters via CLI, dashboard, or API.
  • Distributed Task Schedulers - Allocates compute resources and manages job queues to execute parallel workloads across multiple nodes or clusters.
  • GPU Cluster Job Schedulers - Defines specific rules for how deep learning operations are submitted and scheduled across a GPU compute cluster.
  • Kubernetes ML Platforms - Provides a Kubernetes-native control plane for managing distributed deep learning workloads and ML pipelines.
  • MPI Cluster Orchestrators - Provides MPI cluster orchestration to coordinate large-scale model training across multiple distributed nodes.
  • Lifecycle Orchestration - Manages the complete lifecycle of ML applications, from container building and training to performance monitoring.
  • Native Kubernetes Resource Integration - Leverages native Kubernetes API resources to manage underlying infrastructure and container orchestration.
  • Cluster Resource Managers - Allocates and scales CPU, GPU, and TPU resources across clusters for repeatable ML deployments.
  • DAG Workflow Executions - Coordinates interdependent training and evaluation tasks by executing them as a directed acyclic graph.
  • DAG Workflow Pipelines - Coordinates interdependent operations as a container-native directed acyclic graph to automate complex ML workflows.
  • Experiment Result Comparators - Implements tools for aggregating and contrasting performance metrics from multiple evaluation runs to identify successful model iterations.
  • Model Performance Tracking - Tracks training metrics and objective functions to monitor model convergence and identify the best performing weights.
  • Experiment Visualization - Visualizes the relationships between runs, hyperparameters, and metrics to derive insights from experiment results.
  • Hyperband Optimization - Tunes iterative algorithms by sampling configurations and allocating resources to the best-performing candidates.
  • ML Component Registries - Stores and versions reusable project components with integrated access control and governance.
  • Model Lineage Trackers - Maintains provenance records by associating pipeline artifacts with specific model versions to track data and model history.
  • Grid Search - Executes an exhaustive search across a specified set of discrete values to find optimal configurations.
  • Random Hyperparameter Search - Samples unique configurations from a defined search space to optimize parameters via random search.
  • Sequential Model-Based Optimization - Runs sequential model-based searches using multiple algorithms to find the best configuration for a target metric.
  • Bayesian Optimization - Uses probabilistic methods to compute posterior distributions over objective functions for optimal parameter selection.
  • Run Comparisons - Analyzes hyperparameter versions and training data across different runs to identify high-performing models.
  • Shared Component Hubs - Provides a centralized hub for storing and sharing containerized machine learning logic and tasks across the organization.
  • TensorBoard Dashboards - Integrates TensorBoard dashboards for tracking loss, accuracy, and model graphs during neural network training.
  • Team Collaboration Management - Controls user permissions and data access across projects through administrative roles and team management.
  • Instance Groupings into Projects - Groups related data, connections, and execution components into distinct projects.
  • Workload Isolation Namespaces - Groups resources and users into projects or separate deployments to isolate workloads via Kubernetes namespaces.
  • Run Filtering Queries - Implements a structured query syntax to filter and retrieve experiment runs based on parameters, metrics, or tags.
  • Infrastructure CLI Tooling - Provides a command line interface for executing administrative tasks and interacting with the platform API.
  • Interactive Execution Environments - Provides interactive environments for real-time code execution and data analysis using notebooks and dashboards.
  • Containerized - Stores and versions reusable containerized logic in a centralized hub to ensure consistent deployment across environments.
  • Resource Group Access Controls - Organizes users into teams to manage collective access roles and enforce resource quotas.
  • Automated Container Builds - Packages dependencies and runtime environments into containers for consistent deployment on Kubernetes.
  • CLI Workload Management - Interacts with the orchestration platform through a CLI to trigger and monitor ML operations.
  • Parameterized Execution Sweeps - Runs components across multiple parameter combinations to automate hyperparameter sweeps and repetitive tasks.
  • Federated Cluster Scaling - Distributes compute jobs over multiple namespaces and clusters using agents to handle massive scale.
  • Project-Level Environment Definitions - Defines and validates project-level environment variables and execution contexts for ML tasks.
  • Container Registry Integrations - Integrates the platform with container registries to enable automated building and deployment of images.
  • Data Connection Configurations - Defines secure access points to store and retrieve code, models, datasets, and logs.
  • ML Component Versioning - Manages versioned reusable components with configurable visibility and team-level access controls.
  • Priority-Based Job Schedulers - Routes jobs to specific queues with defined priority and concurrency limits.
  • Job Execution Controllers - Provides runtime control over the start, stop, and execution of automated machine learning workloads.
  • Scheduled Job Execution - Executes training jobs or pipelines automatically based on specific events and reports results to the user.
  • Pipeline Component Sharing - Promotes specific pipeline components with typed inputs and outputs for reuse across projects.
  • Unified Job Abstractions - Offers a unified job abstraction to streamline the execution of builds, experiments, and persistent services.
  • Message-Passing Process Coordination - Coordinates distributed training experiments by managing communication and synchronization between master and worker replicas.
  • Generic Resource Authorization - Evaluates permissions to authorize user access to specific system resources, connections, and namespaces.
  • Organization-Based Access Controls - Separates teams into distinct organizations with independent role-based permissions and resource access controls.
  • Role-Based Access Control - Manages granular permissions and team memberships using a role-based access control system to secure project resources.
  • Secret Management Services - Configures secure connections and secret injection to allow workloads to access private data stores and protected APIs.
  • Configuration Presets - Provides reusable configuration presets to standardize settings across multiple machine learning experiments.
  • Managed Workspace Services - Deploys and manages interactive environments, including notebooks and visualization tools, as managed services.
  • Parallel Execution - Manages the simultaneous execution of multiple nodes within a workflow graph using a mapping abstraction.
  • Pipeline Health Monitors - Streams logs and tracks resource utilization in real time to monitor the operational health and status of active jobs.
  • Experiment Tracking Dashboards - Provides a visual interface to track project progress, view results, and collaborate with teams in real-time.
  • Deep Learning Ecosystems - Platform for reproducible machine learning experiments.
  • Deep Learning Frameworks - Platform for reproducible and scalable machine learning.
  • Machine Learning - Web-based platform for scalable machine learning experiment management.
  • Frameworks de Machine Learning - Platform for reproducible and scalable machine learning experiments.
  • Machine Learning Operations - Web-based platform for reproducible machine learning experiment management.
  • Machine Learning Platform - Platform for reproducible ML on Kubernetes.
  • MLOps Platforms - Orchestrates and manages machine learning platforms.
  • Distributed Frameworks - Platform for reproducible and scalable machine learning workflows.
  • Experiment and Data Management - Platform for reproducible ML workflows on Kubernetes.
  • Infrastructure and Deployment - Lifecycle management for machine learning applications.
  • Gestión de infraestructura - Gestión del ciclo de vida para aplicaciones de aprendizaje automático y aprendizaje profundo.
  • MLOps and Deployment - Platform for scalable machine learning and deep learning.
  • MLOps and Infrastructure - Platform for reproducible and scalable machine learning workflows.
  • MLOps and Lifecycle - MLOps platform.
  • MLOps and Pipelines - Scalable platform for machine learning experiments.
  • Workflow Platforms - Platform for machine learning experimentation and workflows.

Historial de estrellas

Gráfico del historial de estrellas de polyaxon/polyaxonGráfico del historial de estrellas de polyaxon/polyaxon

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Start searching with AI

Alternativas open-source a Polyaxon

Proyectos open-source similares, clasificados según cuántas características comparten con Polyaxon.
  • clearml/clearmlAvatar de clearml

    clearml/clearml

    6,740Ver en GitHub↗

    ClearML is a comprehensive MLOps platform designed to manage the end-to-end machine learning lifecycle, from initial experimentation to production deployment. It provides a suite of integrated tools including a pipeline orchestrator for automating workflows, an experiment tracking tool for logging hyperparameters and metrics, and a metadata-driven data versioning system for managing large-scale datasets and model artifacts. The platform is distinguished by its advanced compute management and serving capabilities. It features a GPU compute manager that supports fractional resource slicing and

    Python
    Ver en GitHub↗6,740
  • transformerlab/transformerlab-appAvatar de transformerlab

    transformerlab/transformerlab-app

    5,103Ver en GitHub↗

    TransformerLab is an MLOps orchestration platform and research environment designed for the training, fine-tuning, and evaluation of large language models. It serves as a centralized control plane for managing machine learning jobs and coordinating distributed GPU compute across hybrid cloud and on-premise providers. The platform distinguishes itself through agent-driven model optimization, using AI assistants to analyze metrics and automatically propose and queue hyperparameter experiments. It provides a remote development environment that allows users to launch interactive notebooks, code e

    Python
    Ver en GitHub↗5,103
  • maiot-io/zenmlAvatar de maiot-io

    maiot-io/zenml

    5,452Ver en GitHub↗

    ZenML is an extensible machine learning orchestration framework designed to manage the end-to-end lifecycle of data pipelines and AI agent workflows. It functions as a durable orchestrator that executes machine learning tasks as directed acyclic graphs, ensuring that every step is containerized for consistent performance across local, cloud, and hybrid infrastructure. By decoupling pipeline code from underlying compute and storage backends, the platform allows developers to define infrastructure-agnostic stacks that remain portable across diverse environments. The project distinguishes itself

    Python
    Ver en GitHub↗5,452
  • tencentmusic/cube-studioAvatar de tencentmusic

    tencentmusic/cube-studio

    5,062Ver en GitHub↗

    Cube Studio is a cloud-native MLOps platform and Kubernetes-based AI orchestrator designed for the entire machine learning lifecycle. It provides a distributed training framework for large-scale model fine-tuning, a GPU resource manager for hardware virtualization, and an ML pipeline orchestrator that uses visual directed acyclic graphs to manage end-to-end workflows. The platform distinguishes itself through its specialized LLM inference server, which supports retrieval-augmented generation and the construction of private knowledge bases. It features a dedicated system for supervised fine-tu

    Pythonaiaihubargo
    Ver en GitHub↗5,062
Ver las 30 alternativas a Polyaxon→

Preguntas frecuentes

¿Qué hace polyaxon/polyaxon?

Polyaxon is a Kubernetes-native machine learning orchestration platform and MLOps pipeline orchestrator. It serves as a control plane for managing distributed deep learning workloads, automated machine learning pipelines, and experiment tracking.

¿Cuáles son las características principales de polyaxon/polyaxon?

Las características principales de polyaxon/polyaxon son: Machine Learning Orchestration, Distributed Training, Centralized Control Planes, Kubernetes Orchestrators, Distributed ML Workload Deployment, AI Workload Orchestration, Data Lineage, Distributed Training Managers.

¿Qué alternativas de código abierto existen para polyaxon/polyaxon?

Las alternativas de código abierto para polyaxon/polyaxon incluyen: clearml/clearml — ClearML is a comprehensive MLOps platform designed to manage the end-to-end machine learning lifecycle, from initial… transformerlab/transformerlab-app — TransformerLab is an MLOps orchestration platform and research environment designed for the training, fine-tuning, and… maiot-io/zenml — ZenML is an extensible machine learning orchestration framework designed to manage the end-to-end lifecycle of data… tencentmusic/cube-studio — Cube Studio is a cloud-native MLOps platform and Kubernetes-based AI orchestrator designed for the entire machine… zenml-io/zenml — ZenML is an orchestration platform designed for building, deploying, and monitoring reproducible machine learning… aws/amazon-sagemaker-examples — This repository is a collection of Jupyter notebooks providing reference implementations and templates for building,…