awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

5 repository-uri

Awesome GitHub RepositoriesDistributed Training Platforms

Systems designed for scaling model training across distributed hardware clusters.

Distinct from Machine Learning Platforms: Distinct from general machine learning platforms: focuses on the distributed training and scaling aspect.

Explore 5 awesome GitHub repositories matching artificial intelligence & ml · Distributed Training Platforms. Refine with filters or upvote what's useful.

Awesome Distributed Training Platforms GitHub Repositories

Găsește cele mai bune repo-uri cu AI.Vom căuta cele mai potrivite repository-uri folosind AI.
  • microsoft/lightgbmAvatar microsoft

    microsoft/LightGBM

    18,096Vezi pe GitHub↗

    LightGBM is a high-performance machine learning framework designed for constructing gradient-boosted decision tree ensembles. It provides a platform for training classification, regression, and ranking models, with a focus on memory efficiency and large-scale distributed computing. The framework distinguishes itself through specialized algorithmic strategies, including leaf-wise tree growth and histogram-based decision learning, which prioritize convergence speed. It optimizes memory usage by bundling mutually exclusive features and employs gradient-based sampling to reduce training complexit

    Scales model training across multiple nodes and hardware accelerators to handle large-scale datasets.

    C++data-miningdecision-treesdistributed
    Vezi pe GitHub↗18,096
  • h2oai/h2o-3Avatar h2oai

    h2oai/h2o-3

    7,493Vezi pe GitHub↗

    h2o-3 is a distributed machine learning platform and automated machine learning framework designed for training and deploying predictive models using distributed in-memory computing. It functions as a deep learning framework and a distributed model scoring engine, capable of operating as a Kubernetes ML cluster to process large datasets in parallel. The platform distinguishes itself through automated machine learning capabilities that automatically select the best algorithms and hyperparameters to optimize model performance. It provides specialized deep learning toolkits for tasks including i

    Provides a platform for scaling machine learning model training across distributed hardware clusters to handle large-scale datasets.

    Jupyter Notebookautomlbig-datadata-science
    Vezi pe GitHub↗7,493
  • alibaba/x-deeplearningAvatar alibaba

    alibaba/x-deeplearning

    4,301Vezi pe GitHub↗

    This project is a distributed machine learning platform and sparse deep learning framework designed for training and serving models with high-dimensional sparse data. It functions as an online model serving infrastructure and recommendation system engine, enabling real-time item retrieval and scoring using deep tree matching and neural networks. The system distinguishes itself through a multi-task learning framework that optimizes multiple objective functions within a shared representation space. It features a specialized online serving infrastructure that supports dynamic model hot-loading a

    Provides a system for scaling model training across distributed hardware clusters using centralized scheduling.

    PureBasic
    Vezi pe GitHub↗4,301
  • azure/machinelearningnotebooksAvatar Azure

    Azure/MachineLearningNotebooks

    4,354Vezi pe GitHub↗

    Azure Machine Learning Notebooks is a cloud-based environment for developing and executing interactive Jupyter notebooks within a managed machine learning workspace. It provides managed machine learning compute through cloud-based workstations and containerized environments pre-configured with GPU drivers and kernels for high-performance model training. The project functions as a distributed GPU training platform and an ML experiment tracking system to monitor training metrics and version data assets. It also serves as an MLOps pipeline orchestrator for automating modular workflows and a mode

    Provides a platform for executing containerized training jobs across GPU clusters with managed compute resources.

    Jupyter Notebookazureazure-machine-learningazure-ml
    Vezi pe GitHub↗4,354
  • fedml-ai/fedmlAvatar FedML-AI

    FedML-AI/FedML

    4,048Vezi pe GitHub↗

    FedML este o bibliotecă distribuită de antrenare a modelelor de machine learning, un framework de învățare federată și un orchestrator de sarcini GPU. Acesta oferă componentele de sistem de bază necesare pentru a executa antrenarea și fine-tuning-ul modelelor la scară largă pe clustere GPU multi-cloud, on-premise și descentralizate, oferind în același timp un motor dedicat pentru servirea scalabilă a modelelor și un manager de pipeline MLOps pentru gestionarea întregului ciclu de viață. Platforma se distinge prin activarea învățării federate cu conservarea confidențialității pe dispozitive edge descentralizate și silozuri organizaționale, păstrând datele brute pe hardware-ul local. De asemenea, dispune de o piață de calcul cu partajare a resurselor care permite utilizatorilor să contribuie cu capacitate GPU neutilizată la un pool partajat pentru execuția sarcinilor distribuite. Sistemul acoperă o gamă largă de capabilități, inclusiv orchestrarea GPU multi-cloud, gestionarea automatizată a pipeline-urilor de machine learning și implementarea AI la nivel de edge pentru dispozitive IoT și smartphone-uri. De asemenea, integrează instrumente pentru fine-tuning-ul modelelor fundamentale, implementarea inferenței cu latență scăzută și monitorizarea experimentelor de antrenare cu profilarea performanței hardware. Utilizatorii pot lansa și programa sarcini folosind o interfață în linie de comandă și fișiere de configurare declarative.

    Implements a platform designed for scaling model training across distributed hardware clusters, including on-premise and cloud environments.

    Python
    Vezi pe GitHub↗4,048
  1. Home
  2. Artificial Intelligence & ML
  3. Machine Learning
  4. Infrastructure
  5. Integrated Development Platforms
  6. Machine Learning Platforms
  7. Distributed Training Platforms