11 रिपॉजिटरी
Methods that partition large-scale model tensors across multiple compute nodes to facilitate parallel processing.
Explore 11 awesome GitHub repositories matching networking & communication · Distributed Parameter Sharding. Refine with filters or upvote what's useful.
TensorFlow is a comprehensive machine learning framework designed for the construction, training, and deployment of complex mathematical models. It utilizes a graph-based execution model that represents operations as directed acyclic graphs, enabling automatic differentiation and efficient parallel processing. The system provides high-level interfaces for defining neural network architectures, alongside a robust engine for managing multidimensional array structures and tensor mathematics. The framework distinguishes itself through a scalable distributed runtime that orchestrates workloads acr
Partitions large-scale model tensors across multiple compute nodes to streamline parallel training and memory management.
Fairseq is a PyTorch toolkit for sequence-to-sequence modeling, specializing in neural machine translation, automatic speech recognition, and large-scale language model training. It provides a framework for processing and aligning diverse data sources, including text, audio, and video, to support tasks such as speech-to-text conversion and multimodal sequence learning. The project is distinguished by its distributed training capabilities, which utilize parameter sharding, mixed-precision training, and CPU offloading to handle models that exceed single-device memory. It also includes specializ
Distributes model parameters and optimizer states across multiple devices to bypass single-GPU memory limits.
Llama 3 is a collection of pretrained, autoregressive transformer-based models designed for natural language generation, reasoning, and complex instruction following. It functions as a generative AI framework that provides the infrastructure for managing model weights, executing neural network inference, and handling computational workloads across diverse knowledge domains. The project distinguishes itself through an integrated AI safety toolkit that employs secondary classification filtering to inspect inputs and outputs, ensuring adherence to usage compliance and safety standards. It suppor
Supports distributed model sharding to partition large neural network parameters across multiple hardware devices.
This project is a machine learning array framework and tensor computation library designed for high-performance numerical computing. It provides a comprehensive suite of tools for constructing and training neural networks, featuring an automatic differentiation engine that facilitates gradient-based optimization and complex mathematical modeling. The library distinguishes itself through a unified memory architecture that allows data to be shared across CPU and GPU devices without explicit copies, significantly reducing data movement overhead. Its execution model relies on a lazy evaluation en
Splits model parameters across multiple devices in-place to reduce memory footprint.
Mamba is a deep learning framework designed for building and training sequence models that process long-range data dependencies with linear-time computational efficiency. By utilizing selective state space modeling, the library enables the construction of neural network architectures that replace traditional attention mechanisms with high-performance state space operations. The framework distinguishes itself through the use of data-dependent state gating, which allows the model to dynamically filter information flow based on the input sequence. To ensure high throughput, it incorporates hardw
Splits model parameters and sequence processing across multiple devices using tensor parallelism.
Implements distributed parameter sharding to partition model tensors across multiple GPUs.
Flax is a deep learning framework and JAX neural network library designed for building complex machine learning models. It functions as a distributed training library and model state manager, providing a toolkit for defining flexible neural network architectures and scaling their training across multiple hardware devices. The project is characterized by a design that separates network logic from parameter values to remain compatible with pure functions. It uses hierarchical module composition to organize networks as trees of nested modules and employs a reference-based state management system
Implements techniques to partition large-scale model tensors across multiple hardware accelerators for parallel processing.
Gemma ओपन-वेट्स लार्ज लैंग्वेज मॉडल का एक परिवार है जो डिकोडर-ओनली ट्रांसफार्मर आर्किटेक्चर पर आधारित है। ये मॉडल टेक्स्ट जनरेशन और मल्टी-मॉडल बातचीत के लिए डिज़ाइन किए गए हैं, जो टेक्स्टुअल और विज़ुअल इनपुट सीक्वेंस दोनों के आधार पर प्रतिक्रियाओं को संसाधित करने और उत्पन्न करने में सक्षम हैं। यह प्रोजेक्ट एक फाइन-ट्यून करने योग्य AI मॉडल प्रदान करता है जो विशेष कार्यों के लिए प्रदर्शन को विशिष्ट बनाने के लिए वेट एडजस्टमेंट और लो-रैंक एडेप्टेशन का समर्थन करता है। इसमें सीमित हार्डवेयर पर मेमोरी उपयोग को कम करने और इन्फरेंस गति बढ़ाने के लिए क्वांटाइज़्ड वेट्स का समर्थन शामिल है। क्षमता सतह मल्टी-मॉडल AI एकीकरण, पैरामीटर शार्डिंग के माध्यम से मेमोरी ऑप्टिमाइज़ेशन, और वास्तविक समय डेटा प्राप्त करने के लिए बाहरी टूल और API के एकीकरण को कवर करती है। यह टेक्स्ट से छवियों के निर्माण और संरचित टेक्स्ट आउटपुट के नमूने को भी सक्षम बनाता है।
Partitions large-scale model tensors across multiple compute nodes to facilitate parallel processing and overcome memory limits.
trlx एक रीइन्फोर्समेंट लर्निंग लाइब्रेरी और ट्रेनिंग फ्रेमवर्क है जिसे मानव फीडबैक का उपयोग करके बड़े भाषा मॉडल्स को संरेखित (align) करने के लिए डिज़ाइन किया गया है। यह कई GPUs और नोड्स में उच्च-पैरामीटर मॉडल्स को स्केल करने के लिए एक डिस्ट्रीब्यूटेड ट्रेनर और कंप्यूट ऑर्केस्ट्रेटर के रूप में कार्य करता है। यह प्रोजेक्ट मानव फीडबैक से रीइन्फोर्समेंट लर्निंग और मॉडल अलाइनमेंट के लिए टूल्स प्रदान करता है। यह लक्ष्य-उन्मुख पुरस्कारों या मानव-लेबल वाले डेटासेट के आधार पर मॉडल व्यवहार को परिष्कृत करने के लिए रिवॉर्ड-मॉडल-आधारित ऑप्टिमाइज़ेशन और प्रॉक्सिमल पॉलिसी ऑप्टिमाइज़ेशन को लागू करता है। यह फ्रेमवर्क मॉडल पैरेललिज़्म, पैरामीटर शार्डिंग और मल्टी-नोड ग्रेडिएंट सिंक्रोनाइज़ेशन सहित डिस्ट्रीब्यूटेड ट्रेनिंग रणनीतियों को कवर करता है। यह रीइन्फोर्समेंट लर्निंग प्रक्रिया के दौरान मॉडल ड्रिफ्ट को मैनेज करने के लिए KL-डाइवर्जेंस जैसे बाधाओं को भी शामिल करता है।
Partitions large model tensors across multiple compute nodes to fit high-parameter architectures in memory.
This project is a distributed machine learning platform and sparse deep learning framework designed for training and serving models with high-dimensional sparse data. It functions as an online model serving infrastructure and recommendation system engine, enabling real-time item retrieval and scoring using deep tree matching and neural networks. The system distinguishes itself through a multi-task learning framework that optimizes multiple objective functions within a shared representation space. It features a specialized online serving infrastructure that supports dynamic model hot-loading a
Allocates parameters globally and merges requests to eliminate communication hotspots in parameter servers.
Higgsfield is a distributed machine learning training framework and GPU cluster orchestrator designed for scaling neural networks with billions of parameters. It functions as a large model sharding system and a containerized deployment tool to manage computational workflows across heterogeneous compute resources. The platform provides a centralized interface for experiment management, enabling the monitoring of real-time telemetry, performance metrics, and logs. It ensures reproducible results by using container isolation to standardize dependencies across different computing environments. T
Implements distributed parameter sharding to train neural networks that exceed the memory capacity of a single GPU.