awesome-repositories.com
Blog
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektÜber unsRanking-MethodikPresseMCP-Server
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
FedML-AI avatar

FedML-AI/FedML

0
View on GitHub↗
4,048 Stars·767 Forks·Python·Apache-2.0·9 AufrufeTensorOpera.ai↗

FedML

FedML ist eine Bibliothek für verteiltes Machine Learning-Training, ein Framework für Federated Learning und ein Orchestrator für GPU-Workloads. Es bietet die Kernsystemkomponenten, die für die Ausführung von groß angelegtem Modelltraining und Fine-Tuning über Multi-Cloud-, On-Premise- und dezentrale GPU-Cluster hinweg erforderlich sind, und bietet zudem eine dedizierte Engine für skalierbares Model-Serving sowie einen MLOps-Pipeline-Manager für das End-to-End-Lifecycle-Management.

Die Plattform zeichnet sich dadurch aus, dass sie datenschutzfreundliches Federated Learning über dezentrale Edge-Geräte und organisatorische Silos hinweg ermöglicht, wobei Rohdaten auf der lokalen Hardware verbleiben. Sie bietet zudem einen Compute-Marktplatz für Ressourcen-Pooling, der es Benutzern ermöglicht, ungenutzte GPU-Kapazitäten für die verteilte Aufgabenausführung in einen gemeinsamen Pool einzubringen.

Das System deckt ein breites Spektrum an Funktionen ab, darunter Multi-Cloud-GPU-Orchestrierung, automatisiertes Machine-Learning-Pipeline-Management und Edge-AI-Deployment für IoT-Geräte und Smartphones. Zudem integriert es Tools für das Fine-Tuning von Foundation-Modellen, Deployment von Inferenz mit geringer Latenz und das Tracking von Trainingsexperimenten mit Hardware-Performance-Profiling.

Benutzer können Workloads über eine Command-Line-Interface und deklarative Konfigurationsdateien starten und planen.

Features

  • Distributed Training Frameworks - Provides a distributed training framework for executing large-scale model training and fine-tuning across multi-cloud and decentralized GPU clusters.
  • AI Workload Orchestration - Pairs AI jobs with economical GPU resources and auto-provisions environments to eliminate manual setup.
  • Distributed GPU Training - Provides a system for distributing the computational load of large-scale model training across multiple GPUs and multi-cloud environments.
  • Distributed Training Orchestration - Coordinates synchronization and parallelization across compute clusters for large-scale model training.
  • LLM Fine-Tuning - Adapts large language models using serverless workflows and pre-built training templates.
  • Large-Scale Model Training - Executes distributed machine learning workloads across clusters to train foundational models exceeding single-device capacity.
  • Distributed Training Platforms - Implements a platform designed for scaling model training across distributed hardware clusters, including on-premise and cloud environments.
  • Decentralized Machine Learning - Executes machine learning tasks across distributed users or edge nodes while keeping raw data localized.
  • Distributed Training - Enables the scaling of machine learning model training and fine-tuning across multiple compute nodes to accelerate workflows.
  • Edge AI Model Deployment - Installs toolkits on smartphones and IoT devices to enable local model training and execution.
  • Model Inference and Serving - Deploys trained models to production environments with a focus on low-latency inference and high-scalability.
  • Federated Learning Frameworks - Offers a unified framework for collaborative machine learning and training across decentralized edge devices and organizational silos.
  • Model Serving & Deployment - Implements infrastructure for hosting and serving machine learning models across various compute planes via network endpoints.
  • Privacy-Preserving Model Training - Trains models across decentralized edge nodes and smartphones while sharing only model updates to maintain data privacy.
  • Foundational Model Adaptation - Adapts open-source foundational models using specific datasets for deployment on GPU clusters.
  • Federated Learning - Trains shared machine learning models across decentralized edge devices and organizational silos without centralizing private data.
  • AI Workload Schedulers - Schedules machine learning workloads across on-premises clusters and multiple GPU cloud environments.
  • GPU Cluster Management Platforms - Manages resource allocation and concurrent job execution via queues for private or on-premises GPU hardware.
  • GPU Resource Orchestrators - Dynamically manages and allocates GPU hardware across shared AI workloads across different cloud providers.
  • GPU Workload Orchestration - Schedules and provisions AI workloads across different cloud providers and private hardware to optimize cost and resource utilization.
  • GPU Cluster Job Schedulers - Provides orchestration for resource allocation and task execution across heterogeneous GPU clusters.
  • End-to-End Training Pipelines - Orchestrates end-to-end machine learning operation pipelines across multi-cloud and edge environments.
  • Distributed Training - Distributes the training of foundational models across decentralized GPUs, multi-cloud environments, and edge servers.
  • Memory-Constrained Inference - Implements specialized memory management and optimization to run large models on hardware with limited VRAM.
  • Experiment Tracking - Records metrics, logs, and hardware performance to compare and evaluate different training runs.
  • Model Serving Engines - Provides an optimized engine for low-latency and scalable inference of large language models on diverse compute planes.
  • Low-Latency Serving Techniques - Provides a framework for deploying models using techniques that minimize inference latency and maximize throughput.
  • ML Pipeline Automation - Orchestrates sequences of interdependent training jobs and data management tasks into automated machine learning pipelines.
  • GPU Environment Provisioning - Automates the setup of execution environments and GPU resource allocation via configuration files.
  • AI Model Production Deployment - Transitions trained models to cloud or on-premise servers for industrial-level scaling.
  • GPU Training Deployments - Provisions and executes machine learning tasks across multi-cloud or on-premises GPU resources.
  • GPU Resource Allocators - Controls the allocation and placement of workloads across available GPU resources to optimize hardware utilization.
  • AI Job Launchers - Runs machine learning jobs on public, decentralized, or on-premise GPU clusters via a command-line interface.
  • Compute Resource Monetization - Allows users to earn payment by contributing unused GPU capacity to a shared compute marketplace.
  • GPU Resource Provisioning - Automates the allocation of GPU compute resources and the setup of the required execution environment.
  • GPU Resource Pooling - Aggregates GPU hardware from multiple contributors into a shared pool for distributed machine learning execution.
  • GPU Compute Marketplaces - Aggregates idle GPU capacity from contributors into a shared pool for on-demand AI task execution.
  • Training Workflow Orchestrators - Orchestrates sequences of interdependent training jobs into a single automated pipeline.
  • Serverless Training Pipelines - Orchestrates interdependent machine learning jobs into automated serverless pipelines.
  • Job Performance Profiling - Tracks real-time billing, system metrics, and logs to diagnose performance bottlenecks in training jobs.
  • Federated Learning - Library for secure and collaborative decentralized machine learning.
  • Federated Learning Frameworks - A research-oriented library for distributed and federated machine learning.
  • Privacy and Safety - Platform for distributed and federated machine learning.

Star-Verlauf

Star-Verlauf für fedml-ai/fedmlStar-Verlauf für fedml-ai/fedml

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Start searching with AI

Open-Source-Alternativen zu FedML

Ähnliche Open-Source-Projekte, sortiert nach der Anzahl der gemeinsamen Funktionen mit FedML.
  • openmlsys/openmlsysAvatar von openmlsys

    openmlsys/openmlsys

    4,813Auf GitHub ansehen↗

    This project is a comprehensive educational resource and curriculum focused on the design and implementation of the full machine learning software and hardware stack. It serves as a technical reference for architecting machine learning systems, spanning from low-level programming interfaces to large-scale deployment infrastructure. The project provides instructional guidance on several specialized domains, including the development of AI compilers through intermediate representations and graph optimizations. It covers the architectural patterns required for distributed training across GPU clu

    TeXcomputer-systemsmachine-learningsoftware-architecture
    Auf GitHub ansehen↗4,813
  • tencentmusic/cube-studioAvatar von tencentmusic

    tencentmusic/cube-studio

    5,062Auf GitHub ansehen↗

    Cube Studio is a cloud-native MLOps platform and Kubernetes-based AI orchestrator designed for the entire machine learning lifecycle. It provides a distributed training framework for large-scale model fine-tuning, a GPU resource manager for hardware virtualization, and an ML pipeline orchestrator that uses visual directed acyclic graphs to manage end-to-end workflows. The platform distinguishes itself through its specialized LLM inference server, which supports retrieval-augmented generation and the construction of private knowledge bases. It features a dedicated system for supervised fine-tu

    Pythonaiaihubargo
    Auf GitHub ansehen↗5,062
  • snowkylin/tensorflow-handbookAvatar von snowkylin

    snowkylin/tensorflow-handbook

    3,927Auf GitHub ansehen↗

    This project is a comprehensive educational resource and tutorial handbook for building, training, and deploying machine learning models using TensorFlow 2. It serves as a structured learning guide covering core deep learning concepts, including neural network architectures, automatic differentiation, and tensor operations. The handbook provides technical guidance on optimizing execution efficiency through GPU memory management, distributed training, and model quantization. It also includes detailed manuals for constructing high-performance data pipelines and exporting models for production s

    Jupyter Notebook
    Auf GitHub ansehen↗3,927
  • infrasys-ai/aisystemAvatar von Infrasys-AI

    Infrasys-AI/AISystem

    17,017Auf GitHub ansehen↗

    AISystem is a comprehensive AI full-stack infrastructure project covering the entire pipeline from AI chip architecture to high-level training frameworks. It encompasses the development of AI compiler frameworks, inference engines, and distributed training orchestrators designed to coordinate workloads across a heterogeneous compute stack of CPUs, GPUs, and NPUs. The project focuses on the deep integration of software and hardware, employing software-hardware co-design to align tensor layouts with physical memory structures. It provides specialized capabilities for accelerating Transformer mo

    Jupyter Notebookaiaiinfraaisys
    Auf GitHub ansehen↗17,017
Alle 30 Alternativen zu FedML anzeigen→

Häufig gestellte Fragen

Was macht fedml-ai/fedml?

FedML ist eine Bibliothek für verteiltes Machine Learning-Training, ein Framework für Federated Learning und ein Orchestrator für GPU-Workloads. Es bietet die Kernsystemkomponenten, die für die Ausführung von groß angelegtem Modelltraining und Fine-Tuning über Multi-Cloud-, On-Premise- und dezentrale GPU-Cluster hinweg erforderlich sind, und bietet zudem eine dedizierte Engine für skalierbares Model-Serving sowie einen MLOps-Pipeline-Manager für das…

Was sind die Hauptfunktionen von fedml-ai/fedml?

Die Hauptfunktionen von fedml-ai/fedml sind: Distributed Training Frameworks, AI Workload Orchestration, Distributed GPU Training, Distributed Training Orchestration, LLM Fine-Tuning, Large-Scale Model Training, Distributed Training Platforms, Decentralized Machine Learning.

Welche Open-Source-Alternativen gibt es zu fedml-ai/fedml?

Open-Source-Alternativen zu fedml-ai/fedml sind unter anderem: openmlsys/openmlsys — This project is a comprehensive educational resource and curriculum focused on the design and implementation of the… tencentmusic/cube-studio — Cube Studio is a cloud-native MLOps platform and Kubernetes-based AI orchestrator designed for the entire machine… snowkylin/tensorflow-handbook — This project is a comprehensive educational resource and tutorial handbook for building, training, and deploying… infrasys-ai/aisystem — AISystem is a comprehensive AI full-stack infrastructure project covering the entire pipeline from AI chip… clearml/clearml — ClearML is a comprehensive MLOps platform designed to manage the end-to-end machine learning lifecycle, from initial… maiot-io/zenml — ZenML is an extensible machine learning orchestration framework designed to manage the end-to-end lifecycle of data…