awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

5 个仓库

Awesome GitHub RepositoriesDistributed Training Platforms

Systems designed for scaling model training across distributed hardware clusters.

Distinct from Machine Learning Platforms: Distinct from general machine learning platforms: focuses on the distributed training and scaling aspect.

Explore 5 awesome GitHub repositories matching artificial intelligence & ml · Distributed Training Platforms. Refine with filters or upvote what's useful.

Awesome Distributed Training Platforms GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • microsoft/lightgbmmicrosoft 的头像

    microsoft/LightGBM

    18,096在 GitHub 上查看↗

    LightGBM is a high-performance machine learning framework designed for constructing gradient-boosted decision tree ensembles. It provides a platform for training classification, regression, and ranking models, with a focus on memory efficiency and large-scale distributed computing. The framework distinguishes itself through specialized algorithmic strategies, including leaf-wise tree growth and histogram-based decision learning, which prioritize convergence speed. It optimizes memory usage by bundling mutually exclusive features and employs gradient-based sampling to reduce training complexit

    Scales model training across multiple nodes and hardware accelerators to handle large-scale datasets.

    C++data-miningdecision-treesdistributed
    在 GitHub 上查看↗18,096
  • h2oai/h2o-3h2oai 的头像

    h2oai/h2o-3

    7,493在 GitHub 上查看↗

    h2o-3 is a distributed machine learning platform and automated machine learning framework designed for training and deploying predictive models using distributed in-memory computing. It functions as a deep learning framework and a distributed model scoring engine, capable of operating as a Kubernetes ML cluster to process large datasets in parallel. The platform distinguishes itself through automated machine learning capabilities that automatically select the best algorithms and hyperparameters to optimize model performance. It provides specialized deep learning toolkits for tasks including i

    Provides a platform for scaling machine learning model training across distributed hardware clusters to handle large-scale datasets.

    Jupyter Notebookautomlbig-datadata-science
    在 GitHub 上查看↗7,493
  • alibaba/x-deeplearningalibaba 的头像

    alibaba/x-deeplearning

    4,301在 GitHub 上查看↗

    This project is a distributed machine learning platform and sparse deep learning framework designed for training and serving models with high-dimensional sparse data. It functions as an online model serving infrastructure and recommendation system engine, enabling real-time item retrieval and scoring using deep tree matching and neural networks. The system distinguishes itself through a multi-task learning framework that optimizes multiple objective functions within a shared representation space. It features a specialized online serving infrastructure that supports dynamic model hot-loading a

    Provides a system for scaling model training across distributed hardware clusters using centralized scheduling.

    PureBasic
    在 GitHub 上查看↗4,301
  • azure/machinelearningnotebooksAzure 的头像

    Azure/MachineLearningNotebooks

    4,354在 GitHub 上查看↗

    Azure Machine Learning Notebooks is a cloud-based environment for developing and executing interactive Jupyter notebooks within a managed machine learning workspace. It provides managed machine learning compute through cloud-based workstations and containerized environments pre-configured with GPU drivers and kernels for high-performance model training. The project functions as a distributed GPU training platform and an ML experiment tracking system to monitor training metrics and version data assets. It also serves as an MLOps pipeline orchestrator for automating modular workflows and a mode

    Provides a platform for executing containerized training jobs across GPU clusters with managed compute resources.

    Jupyter Notebookazureazure-machine-learningazure-ml
    在 GitHub 上查看↗4,354
  • fedml-ai/fedmlFedML-AI 的头像

    FedML-AI/FedML

    4,048在 GitHub 上查看↗

    FedML 是一个分布式机器学习训练库、联邦学习框架和 GPU 工作负载编排器。它提供了在多云、本地和去中心化 GPU 集群上执行大规模模型训练和微调所需的核心系统组件,同时为可扩展的模型服务提供专用引擎,并为端到端生命周期管理提供 MLOps 流水线管理器。 该平台的独特之处在于支持跨去中心化边缘设备和组织孤岛的隐私保护联邦学习,将原始数据保留在本地硬件上。它还具有资源池化计算市场,允许用户将未使用的 GPU 容量贡献给共享池以进行分布式任务执行。 该系统涵盖了广泛的功能,包括多云 GPU 编排、自动化机器学习流水线管理以及针对物联网设备和智能手机的边缘 AI 部署。它进一步集成了用于基础模型微调、低延迟推理部署以及带有硬件性能分析的训练实验追踪工具。 用户可以使用命令行界面和声明式配置文件来启动和调度工作负载。

    Implements a platform designed for scaling model training across distributed hardware clusters, including on-premise and cloud environments.

    Python
    在 GitHub 上查看↗4,048
  1. Home
  2. Artificial Intelligence & ML
  3. Machine Learning
  4. Infrastructure
  5. Integrated Development Platforms
  6. Machine Learning Platforms
  7. Distributed Training Platforms