awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
FedML-AI avatar

FedML-AI/FedML

0
View on GitHub↗
4,048 stars·767 forks·Python·Apache-2.0·26 viewsTensorOpera.ai↗

FedML

FedML is a distributed machine learning training library, federated learning framework, and GPU workload orchestrator. It provides the core system components necessary to execute large-scale model training and fine-tuning across multi-cloud, on-premise, and decentralized GPU clusters, while offering a dedicated engine for scalable model serving and an MLOps pipeline manager for end-to-end lifecycle management.

The platform distinguishes itself by enabling privacy-preserving federated learning across decentralized edge devices and organizational silos, keeping raw data on local hardware. It also features a resource-pooling compute marketplace that allows users to contribute unused GPU capacity to a shared pool for distributed task execution.

The system covers a broad range of capabilities, including multi-cloud GPU orchestration, automated machine learning pipeline management, and edge AI deployment for IoT devices and smartphones. It further integrates tools for foundational model fine-tuning, low-latency inference deployment, and training experiment tracking with hardware performance profiling.

Users can launch and schedule workloads using a command-line interface and declarative configuration files.

Features

  • Distributed Training Frameworks - Provides a distributed training framework for executing large-scale model training and fine-tuning across multi-cloud and decentralized GPU clusters.
  • AI Workload Orchestration - Pairs AI jobs with economical GPU resources and auto-provisions environments to eliminate manual setup.
  • Distributed GPU Training - Provides a system for distributing the computational load of large-scale model training across multiple GPUs and multi-cloud environments.
  • Distributed Training Orchestration - Coordinates synchronization and parallelization across compute clusters for large-scale model training.
  • LLM Fine-Tuning - Adapts large language models using serverless workflows and pre-built training templates.
  • Large-Scale Model Training - Executes distributed machine learning workloads across clusters to train foundational models exceeding single-device capacity.
  • Distributed Training Platforms - Implements a platform designed for scaling model training across distributed hardware clusters, including on-premise and cloud environments.
  • Decentralized Machine Learning - Executes machine learning tasks across distributed users or edge nodes while keeping raw data localized.
  • Distributed Training - Enables the scaling of machine learning model training and fine-tuning across multiple compute nodes to accelerate workflows.
  • Edge AI Model Deployment - Installs toolkits on smartphones and IoT devices to enable local model training and execution.
  • Model Inference and Serving - Deploys trained models to production environments with a focus on low-latency inference and high-scalability.
  • Federated Learning Frameworks - Offers a unified framework for collaborative machine learning and training across decentralized edge devices and organizational silos.
  • Model Serving & Deployment - Implements infrastructure for hosting and serving machine learning models across various compute planes via network endpoints.
  • Privacy-Preserving Model Training - Trains models across decentralized edge nodes and smartphones while sharing only model updates to maintain data privacy.
  • Foundational Model Adaptation - Adapts open-source foundational models using specific datasets for deployment on GPU clusters.
  • Federated Learning - Trains shared machine learning models across decentralized edge devices and organizational silos without centralizing private data.
  • AI Workload Schedulers - Schedules machine learning workloads across on-premises clusters and multiple GPU cloud environments.
  • GPU Cluster Management Platforms - Manages resource allocation and concurrent job execution via queues for private or on-premises GPU hardware.
  • GPU Resource Orchestrators - Dynamically manages and allocates GPU hardware across shared AI workloads across different cloud providers.
  • GPU Workload Orchestration - Schedules and provisions AI workloads across different cloud providers and private hardware to optimize cost and resource utilization.
  • GPU Cluster Job Schedulers - Provides orchestration for resource allocation and task execution across heterogeneous GPU clusters.
  • End-to-End Training Pipelines - Orchestrates end-to-end machine learning operation pipelines across multi-cloud and edge environments.
  • Distributed Training - Distributes the training of foundational models across decentralized GPUs, multi-cloud environments, and edge servers.
  • Memory-Constrained Inference - Implements specialized memory management and optimization to run large models on hardware with limited VRAM.
  • Experiment Tracking - Records metrics, logs, and hardware performance to compare and evaluate different training runs.
  • Model Serving Engines - Provides an optimized engine for low-latency and scalable inference of large language models on diverse compute planes.
  • Low-Latency Serving Techniques - Provides a framework for deploying models using techniques that minimize inference latency and maximize throughput.
  • ML Pipeline Automation - Orchestrates sequences of interdependent training jobs and data management tasks into automated machine learning pipelines.
  • GPU Environment Provisioning - Automates the setup of execution environments and GPU resource allocation via configuration files.
  • AI Model Production Deployment - Transitions trained models to cloud or on-premise servers for industrial-level scaling.
  • GPU Training Deployments - Provisions and executes machine learning tasks across multi-cloud or on-premises GPU resources.
  • GPU Resource Allocators - Controls the allocation and placement of workloads across available GPU resources to optimize hardware utilization.
  • AI Job Launchers - Runs machine learning jobs on public, decentralized, or on-premise GPU clusters via a command-line interface.
  • Compute Resource Monetization - Allows users to earn payment by contributing unused GPU capacity to a shared compute marketplace.
  • GPU Resource Provisioning - Automates the allocation of GPU compute resources and the setup of the required execution environment.
  • GPU Resource Pooling - Aggregates GPU hardware from multiple contributors into a shared pool for distributed machine learning execution.
  • GPU Compute Marketplaces - Aggregates idle GPU capacity from contributors into a shared pool for on-demand AI task execution.
  • Training Workflow Orchestrators - Orchestrates sequences of interdependent training jobs into a single automated pipeline.
  • Serverless Training Pipelines - Orchestrates interdependent machine learning jobs into automated serverless pipelines.
  • Job Performance Profiling - Tracks real-time billing, system metrics, and logs to diagnose performance bottlenecks in training jobs.
  • Federated Learning - Library for secure and collaborative decentralized machine learning.
  • Federated Learning Frameworks - A research-oriented library for distributed and federated machine learning.
  • Privacy and Safety - Platform for distributed and federated machine learning.

Star history

Star history chart for fedml-ai/fedmlStar history chart for fedml-ai/fedml

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with FedML

These projects share indexed features with FedML. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • openmlsys/openmlsysopenmlsys avatar

    openmlsys/openmlsys

    4,813View on GitHub↗

    This project is a comprehensive educational resource and curriculum focused on the design and implementation of the full machine learning software and hardware stack. It serves as a technical reference for architecting machine learning systems, spanning from low-level programming interfaces to large-scale deployment infrastructure. The project provides instructional guidance on several specialized domains, including the development of AI compilers through intermediate representations and graph optimizations. It covers the architectural patterns required for distributed training across GPU clu

    TeXcomputer-systemsmachine-learningsoftware-architecture
    View on GitHub↗4,813
  • tencentmusic/cube-studiotencentmusic avatar

    tencentmusic/cube-studio

    5,062View on GitHub↗

    Cube Studio is a cloud-native MLOps platform and Kubernetes-based AI orchestrator designed for the entire machine learning lifecycle. It provides a distributed training framework for large-scale model fine-tuning, a GPU resource manager for hardware virtualization, and an ML pipeline orchestrator that uses visual directed acyclic graphs to manage end-to-end workflows. The platform distinguishes itself through its specialized LLM inference server, which supports retrieval-augmented generation and the construction of private knowledge bases. It features a dedicated system for supervised fine-tu

    Pythonaiaihubargo
    View on GitHub↗5,062
  • snowkylin/tensorflow-handbooksnowkylin avatar

    snowkylin/tensorflow-handbook

    3,927View on GitHub↗

    This project is a comprehensive educational resource and tutorial handbook for building, training, and deploying machine learning models using TensorFlow 2. It serves as a structured learning guide covering core deep learning concepts, including neural network architectures, automatic differentiation, and tensor operations. The handbook provides technical guidance on optimizing execution efficiency through GPU memory management, distributed training, and model quantization. It also includes detailed manuals for constructing high-performance data pipelines and exporting models for production s

    Jupyter Notebook
    View on GitHub↗3,927
  • infrasys-ai/aisystemInfrasys-AI avatar

    Infrasys-AI/AISystem

    17,017View on GitHub↗

    AISystem is a comprehensive AI full-stack infrastructure project covering the entire pipeline from AI chip architecture to high-level training frameworks. It encompasses the development of AI compiler frameworks, inference engines, and distributed training orchestrators designed to coordinate workloads across a heterogeneous compute stack of CPUs, GPUs, and NPUs. The project focuses on the deep integration of software and hardware, employing software-hardware co-design to align tensor layouts with physical memory structures. It provides specialized capabilities for accelerating Transformer mo

    Jupyter Notebookaiaiinfraaisys
    View on GitHub↗17,017
Compare all 30 related projects→

Frequently asked questions

What does fedml-ai/fedml do?

FedML is a distributed machine learning training library, federated learning framework, and GPU workload orchestrator. It provides the core system components necessary to execute large-scale model training and fine-tuning across multi-cloud, on-premise, and decentralized GPU clusters, while offering a dedicated engine for scalable model serving and an MLOps pipeline manager for end-to-end lifecycle management.

What are the main features of fedml-ai/fedml?

The main features of fedml-ai/fedml are: Distributed Training Frameworks, AI Workload Orchestration, Distributed GPU Training, Distributed Training Orchestration, LLM Fine-Tuning, Large-Scale Model Training, Distributed Training Platforms, Decentralized Machine Learning.

Which projects share features with fedml-ai/fedml?

Projects with overlapping indexed features include: openmlsys/openmlsys — This project is a comprehensive educational resource and curriculum focused on the design and implementation of the… tencentmusic/cube-studio — Cube Studio is a cloud-native MLOps platform and Kubernetes-based AI orchestrator designed for the entire machine… snowkylin/tensorflow-handbook — This project is a comprehensive educational resource and tutorial handbook for building, training, and deploying… infrasys-ai/aisystem — AISystem is a comprehensive AI full-stack infrastructure project covering the entire pipeline from AI chip… clearml/clearml — ClearML is a comprehensive MLOps platform designed to manage the end-to-end machine learning lifecycle, from initial… maiot-io/zenml — ZenML is an extensible machine learning orchestration framework designed to manage the end-to-end lifecycle of data…