awesome-repositories.com
المدونة
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعحولكيفية ترتيب النتائجالصحافةخادم MCP
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
bigscience-workshop avatar

bigscience-workshop/petals

0
View on GitHub↗
10,208 نجوم·615 تفرعات·Python·MIT·6 مشاهداتpetals.dev↗

Petals

Petals is a decentralized framework and inference engine for running large language models across a peer-to-peer network. It enables the execution of models that exceed the memory of any single machine by splitting computations and model layers across a collaborative swarm of GPUs.

The system functions as a collaborative compute network where participants share local GPU resources and host model weights. It supports distributed prompt-tuning to adapt massive models to specific tasks and allows for the establishment of private compute swarms to process sensitive data within restricted, trusted networks.

The platform manages distributed layer execution and pipeline-parallel inference, utilizing distributed hash tables for peer discovery and circuit relays to bypass firewalls. It includes mechanisms for dynamic block hosting and remote weight streaming to optimize how model parameters are loaded and distributed across the swarm.

The software is implemented in Python.

Features

  • Distributed Inference Engines - Provides a framework for splitting and executing massive language model inference across a collaborative network of GPUs.
  • Model Sharding - Splits neural network layers across multiple network nodes to execute models larger than any single machine's memory.
  • Collaborative GPU Swarms - Establishes a decentralized network of connected devices that collectively host model weights and execute inference.
  • Dynamic GPU Block Allocation - Allocates model layers to GPUs based on VRAM and adjusts hosting patterns to optimize swarm throughput.
  • Collaborative GPU Sharing - Enables sharing local GPU resources to serve specific parts of a model and increase collective network capacity.
  • Large Language Model Serving - Implements strategies for hosting and serving models that exceed the memory capacity of any single machine.
  • Model Inference Servers - Provides a distributed inference engine that splits large language model computations across multiple network nodes.
  • Weight Distribution - Implements strategies for splitting and hosting model parameters across multiple GPU devices in a distributed swarm.
  • Pipeline Parallelism Partitioners - Implements pipeline-parallelism by partitioning large neural networks into sequential layers across multiple remote GPUs.
  • Multi-GPU Layer Distribution - Spreads individual model layers across a collaborative network of computers to execute models exceeding single-device memory.
  • ML Model Hosting - Supports loading and serving model weights directly from remote repositories without requiring local format conversion.
  • Peer-to-Peer Networking - Implements a peer-to-peer networking model where nodes collectively host and serve large model weights.
  • Peer Discovery - Utilizes distributed hash tables and bootstrap peers to discover and connect available compute nodes across the internet.
  • Distributed Prompt Tuning - Adapts massive models to specific tasks using prompt-tuning techniques leveraging shared network resources.
  • Large Language Model Fine-Tuning Frameworks - Provides a platform for adapting large models to specific tasks via distributed prompt-tuning.
  • Distributed Fine-Tuning - Enables updating model behavior for specific tasks through distributed prompt-tuning on a network.
  • Private AI Deployments - Enables the creation of private distributed compute networks to ensure data privacy and restricted access.
  • Remote Weight Streaming - Loads model parameters directly from remote repositories into GPU memory to eliminate local file conversion steps.
  • Trusted Host Establishment - Enables establishing restricted swarms of trusted hosts to ensure data is processed only by authorized organizations.
  • Private Networks - Supports launching restricted networks of compute nodes for private model inference and fine-tuning.
  • Circuit Relay Hosting - Implements circuit relaying to route traffic between peers, bypassing firewalls and NAT for nodes without public IPs.
  • Peer Connectivity - Uses bootstrap peers and relays to establish and maintain stable peer-to-peer connectivity across firewalls.
  • Private Data Processing Environments - Enables the creation of restricted networks of trusted hardware to process sensitive data in isolation.
  • Private Network Security - Allows for the creation of restricted, trusted networks to process sensitive data without public swarm exposure.
  • GPU Block Memory Management - Allows controlling which model blocks a server hosts and how many to load based on available GPU memory.
  • Chat Interfaces - Distributed inference and fine-tuning using peer-to-peer resources.
  • Inference Frameworks - Distributed inference and fine-tuning over internet-connected nodes.
  • Large Language Models - Distributed inference and fine-tuning for massive models.
  • Developer Tools and Infrastructure - Distributed platform for running large AI models.
  • Large Language Models (LLMs) - Listed in the “Large Language Models (LLMs)” section of the The Incredible Pytorch awesome list.

سجل النجوم

مخطط تاريخ النجوم لـ bigscience-workshop/petalsمخطط تاريخ النجوم لـ bigscience-workshop/petals

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Start searching with AI

بدائل مفتوحة المصدر لـ Petals

مشاريع مفتوحة المصدر مشابهة، مرتبة حسب عدد الميزات المشتركة مع Petals.
  • lm-sys/fastchatالصورة الرمزية لـ lm-sys

    lm-sys/FastChat

    39,472عرض على GitHub↗

    FastChat is a training and serving platform for large language models that provides an integrated toolkit for fine-tuning, hosting, and benchmarking chatbots. It functions as an inference server capable of hosting multiple models and exposing them via a standardized API for chat applications. The platform distinguishes itself through a distributed model controller that manages worker nodes and routes requests across a hardware-agnostic inference layer supporting various accelerators. It includes a dedicated evaluation framework for assessing model quality using automated judges, multi-turn di

    Python
    عرض على GitHub↗39,472
  • vllm-project/vllmالصورة الرمزية لـ vllm-project

    vllm-project/vllm

    83,048عرض على GitHub↗

    vLLM is a high-throughput inference engine designed for the efficient serving and execution of large language models. It functions as a production-ready distributed model server, providing standard API protocols for online serving while also supporting offline batch processing. The system is built to maximize token generation speed and memory efficiency, enabling both large-scale cloud deployments and local execution on personal hardware. The project distinguishes itself through advanced memory management and request scheduling techniques, most notably its use of non-contiguous key-value cach

    Pythonamdblackwellcuda
    عرض على GitHub↗83,048
  • pytorch/serveالصورة الرمزية لـ pytorch

    pytorch/serve

    4,354عرض على GitHub↗

    This project is a PyTorch model serving framework designed to deploy and scale machine learning models in production via scalable network endpoints. It functions as a high-performance inference server, optimizer, and model lifecycle manager that handles model loading, request batching, and hardware acceleration. The system distinguishes itself through advanced orchestration and optimization capabilities, such as chaining multiple models into sequential workflows using execution graphs and employing dynamic batching to improve throughput and latency. It provides specialized support for generat

    Java
    عرض على GitHub↗4,354
  • yuanzhoulvpi2017/zero_nlpالصورة الرمزية لـ yuanzhoulvpi2017

    yuanzhoulvpi2017/zero_nlp

    3,825عرض على GitHub↗

    zero_nlp is a distributed framework for training and fine-tuning large language models and multimodal architectures. It provides a specialized toolkit for distributed model parallelism, allowing neural network layers and weights to be partitioned across multiple GPU devices to train models that exceed the memory capacity of a single processor. The project distinguishes itself through a combination of high-throughput data pipelines and parameter-efficient tuning. It utilizes multi-threading and memory mapping to preprocess and stream datasets exceeding 100GB and implements memory-saving adapta

    Jupyter Notebookbertchatglm-6bclip
    عرض على GitHub↗3,825
عرض جميع البدائل الـ 30 لـ Petals→

الأسئلة الشائعة

ما هي وظيفة bigscience-workshop/petals؟

Petals is a decentralized framework and inference engine for running large language models across a peer-to-peer network. It enables the execution of models that exceed the memory of any single machine by splitting computations and model layers across a collaborative swarm of GPUs.

ما هي الميزات الرئيسية لـ bigscience-workshop/petals؟

الميزات الرئيسية لـ bigscience-workshop/petals هي: Distributed Inference Engines, Model Sharding, Collaborative GPU Swarms, Dynamic GPU Block Allocation, Collaborative GPU Sharing, Large Language Model Serving, Model Inference Servers, Weight Distribution.

ما هي البدائل مفتوحة المصدر لـ bigscience-workshop/petals؟

تشمل البدائل مفتوحة المصدر لـ bigscience-workshop/petals: lm-sys/fastchat — FastChat is a training and serving platform for large language models that provides an integrated toolkit for… vllm-project/vllm — vLLM is a high-throughput inference engine designed for the efficient serving and execution of large language models.… pytorch/serve — This project is a PyTorch model serving framework designed to deploy and scale machine learning models in production… yuanzhoulvpi2017/zero_nlp — zero_nlp is a distributed framework for training and fine-tuning large language models and multimodal architectures.… i2p/i2p.i2p — I2P is a decentralized anonymous network layer and peer-to-peer overlay network. It functions as a darknet… lightning-ai/litgpt — LitGPT is a training and deployment framework for large language models, providing a suite of tools for pretraining,…