56 रिपॉजिटरी
Libraries for accelerating training, inference, and distributed computing.
Explore 56 awesome GitHub repositories matching part of an awesome list · Computation and Optimization. Refine with filters or upvote what's useful.
TensorFlow is a comprehensive machine learning framework designed for the construction, training, and deployment of complex mathematical models. It utilizes a graph-based execution model that represents operations as directed acyclic graphs, enabling automatic differentiation and efficient parallel processing. The system provides high-level interfaces for defining neural network architectures, alongside a robust engine for managing multidimensional array structures and tensor mathematics. The framework distinguishes itself through a scalable distributed runtime that orchestrates workloads acr
Platform for developing and deploying machine learning applications.
PyTorch is a machine learning framework centered on a GPU-ready tensor library that supports multi-dimensional array operations across both CPU and accelerator hardware. It provides a foundational infrastructure for mathematical computation and dynamic neural network construction, utilizing a tape-based automatic differentiation system that allows for flexible, non-static graph execution. The framework is designed for deep integration with Python, enabling natural usage alongside standard scientific computing ecosystems. It distinguishes itself through a comprehensive distributed training sui
Core library for developing and training deep learning models.
Scikit-learn is a machine learning library for predictive data analysis that provides a collection of algorithms for supervised and unsupervised learning. It functions as a comprehensive toolkit for data preprocessing, dimensionality reduction, and model selection, allowing users to classify data objects, predict continuous values, and cluster similar items based on historical patterns. The project is defined by a unified interface design where objects either learn from data, transform data, or chain these operations into sequential workflows. To ensure performance on large or high-dimensiona
Library for data preparation and statistical model building.
Ray is a distributed computing framework designed to scale Python and Java applications across clusters by abstracting task scheduling and resource management. It functions as a resource-aware execution engine that manages task dependencies, placement, and fault tolerance across networked compute nodes. At its core, the system provides a stateful actor model, allowing developers to define classes that run in dedicated processes to maintain and mutate internal state across remote method calls. The framework distinguishes itself through a robust cross-language interoperability layer, enabling f
Distributed execution framework for machine learning workloads.
DeepSpeed is a high-performance library designed to scale deep learning model training and inference across massive clusters of GPUs and compute nodes. It provides a comprehensive suite of tools for distributed training, enabling the execution of models that exceed the memory capacity of single devices through advanced parameter partitioning, pipeline-based model parallelism, and memory-efficient state offloading. The framework distinguishes itself through specialized communication-efficient optimizers and hardware-aware acceleration techniques. By utilizing gradient compression, quantization
Optimization library for efficient distributed training and inference.
ColossalAI is a distributed deep learning framework designed for training and deploying massive artificial intelligence models across clusters of hardware accelerators. It functions as a parallel computing engine that partitions model workloads and data across multiple processors to maximize memory efficiency and throughput. The platform distinguishes itself through a comprehensive suite of parallelization strategies, including multi-dimensional tensor parallelism and pipeline-based model parallelism, which segment neural network layers and stages across devices. To support large-scale genera
System for efficient large-scale AI model training and inference.
This project is a high-performance numerical computing library designed for large-scale scientific and machine learning workloads. It functions as an automatic differentiation framework and a just-in-time compilation engine, transforming high-level Python code into optimized machine instructions. By enforcing pure functional programming patterns and immutable array semantics, the library ensures that mathematical functions remain compatible with automated graph transformations and symbolic differentiation. The platform distinguishes itself through its distributed array computing capabilities,
Library for composable transformations of Python and NumPy programs.
PyTorch Lightning is a deep learning research framework that provides a structured environment for organizing machine learning code. It functions as a unified trainer orchestrator, centralizing the execution flow by managing the interaction between hardware resources, data loaders, and model components. By decoupling model architecture from training logic, the framework enables researchers to maintain clean, modular codebases that remain portable across different environments. The framework distinguishes itself through a hardware-agnostic abstraction layer that scales deep learning workloads
Interface for training and deploying models on multiple accelerators.
XGBoost is a distributed machine learning library for implementing scalable gradient boosting decision trees used for regression, classification, and ranking. It functions as a predictive model framework and a cross-language toolkit, providing a core implementation with native bindings for Python, R, Java, Scala, and C++. The system is designed as a GPU-accelerated library that utilizes CUDA and NCCL to speed up the training of decision tree ensembles. It operates as a distributed framework capable of scaling training and prediction across multi-node clusters and GPU environments to process m
Optimized distributed gradient boosting library.
This project is a machine learning array framework and tensor computation library designed for high-performance numerical computing. It provides a comprehensive suite of tools for constructing and training neural networks, featuring an automatic differentiation engine that facilitates gradient-based optimization and complex mathematical modeling. The library distinguishes itself through a unified memory architecture that allows data to be shared across CPU and GPU devices without explicit copies, significantly reducing data movement overhead. Its execution model relies on a lazy evaluation en
Array framework optimized for machine learning on Apple silicon.
Paddle is a deep learning framework designed for building, training, and deploying neural networks. It provides a platform for constructing models using tensor-based computations and supports both dynamic and static execution graphs to facilitate research and production workflows. The platform functions as a distributed machine learning system, enabling the scaling of training workloads across multiple nodes and hardware clusters. It includes a comprehensive toolkit for model deployment and optimization, allowing users to convert external model formats, compress trained models for resource-co
Framework for large-scale deep network training across nodes.
This library provides a framework for parameter-efficient fine-tuning, enabling the adaptation of large pretrained models by training only a small subset of parameters. It functions as a distributed model training system and optimization toolkit, designed to reduce the computational and memory requirements typically associated with full model fine-tuning. The project distinguishes itself through a suite of methods for modular adapter composition, including low-rank matrix decomposition and activation-based scaling. It supports the integration of multiple task-specific adapter modules, allowin
Methods for parameter-efficient fine-tuning of pre-trained models.
Triton is a parallel computing framework and high-level programming language designed for writing custom compute kernels. It functions as a deep learning compiler, translating complex mathematical operations into high-throughput instructions that maximize hardware utilization and memory efficiency on graphics processing units. The framework distinguishes itself through a hardware-agnostic compute abstraction that allows developers to define kernels without manual low-level tuning. It employs just-in-time compilation to generate optimized binary instructions at runtime, utilizing static data f
Language and compiler for writing efficient custom deep-learning primitives.
LightGBM is a high-performance machine learning framework designed for constructing gradient-boosted decision tree ensembles. It provides a platform for training classification, regression, and ranking models, with a focus on memory efficiency and large-scale distributed computing. The framework distinguishes itself through specialized algorithmic strategies, including leaf-wise tree growth and histogram-based decision learning, which prioritize convergence speed. It optimizes memory usage by bundling mutually exclusive features and employs gradient-based sampling to reduce training complexit
Gradient boosting framework using tree-based learning algorithms.
Horovod is a distributed deep learning framework and gradient synchronizer designed to scale model training across multiple GPUs and compute nodes. It functions as a distributed training orchestrator and an elastic training engine, utilizing an MPI collective communication library to synchronize weights and gradients across TensorFlow, PyTorch, Keras, and MXNet models. The system distinguishes itself through dynamic elastic scaling, which allows it to adjust the number of active workers at runtime and recover from node failures. It optimizes communication efficiency using tensor fusion batchi
Distributed training framework for TensorFlow, Keras, and PyTorch.
DGL is a Python library for building and training graph neural networks. It functions as a graph message passing framework and a geometric deep learning tool, enabling the development of models that analyze graph-structured data. The library is designed for large-scale graph processing, utilizing distributed training and neighbor sampling to handle datasets with billions of edges. It provides specialized support for heterogeneous graph modeling, allowing for the representation of complex real-world entities with multiple node and edge types. Its capabilities cover a wide range of graph tasks
Scalable Python package for deep learning on graphs.
Dask एक पैरेलल कंप्यूटिंग फ्रेमवर्क और डिस्ट्रीब्यूटेड टास्क शेड्यूलर है जिसे Python डेटा साइंस वर्कफ़्लो को सिंगल मशीनों से बड़े क्लस्टर्स तक स्केल करने के लिए डिज़ाइन किया गया है। यह एक क्लस्टर रिसोर्स मैनेजर के रूप में कार्य करता है जो कार्यों और उनकी डिपेंडेंसी को डायरेक्टेड एसाइक्लिक ग्राफ (DAGs) के रूप में प्रस्तुत करके कम्प्यूटेशनल लॉजिक को व्यवस्थित करता है। यह आर्किटेक्चर सिस्टम को जटिल निष्पादन आवश्यकताओं का प्रबंधन करते हुए उपलब्ध हार्डवेयर पर वर्कलोड के वितरण को स्वचालित करने की अनुमति देता है। यह प्रोजेक्ट एक लेज़ी इवैल्यूएशन इंजन के माध्यम से खुद को अलग करता है जो डेटा ऑपरेशन्स को तब तक स्थगित कर देता है जब तक कि उन्हें स्पष्ट रूप से अनुरोध न किया जाए, जिससे ग्लोबल ग्राफ ऑप्टिमाइज़ेशन और कुशल संसाधन आवंटन सक्षम होता है। इसमें उपलब्ध मेमोरी से अधिक डेटासेट को प्रोसेस करते समय सिस्टम क्रैश को रोकने के लिए मेमोरी-अवेयर डेटा स्पिलिंग शामिल है, और यह टास्क ग्राफ फ्यूजन का उपयोग ऑपरेशन्स के अनुक्रमों को एकल निष्पादन चरणों में संयोजित करने के लिए करता है, जिससे शेड्यूलिंग ओवरहेड और इंटर-नोड संचार कम हो जाता है। यह प्लेटफॉर्म बड़े पैमाने पर डेटा एनालिटिक्स के लिए एक व्यापक क्षमता सतह प्रदान करता है, जिसमें डिस्ट्रीब्यूटेड मशीन लर्निंग, उच्च-प्रदर्शन कंप्यूटिंग एकीकरण, और पैरेलल डेटा प्रोसेसिंग के लिए समर्थन शामिल है। यह क्लस्टर लाइफसाइकिल मैनेजमेंट, परफॉरमेंस प्रोफाइलिंग, और टास्क निष्पादन की रीयल-टाइम मॉनिटरिंग के लिए व्यापक उपकरण प्रदान करता है। उपयोगकर्ता इन वातावरणों को स्थानीय हार्डवेयर, क्लाउड प्रदाताओं, कंटेनरीकृत सिस्टम, और उच्च-प्रदर्शन कंप्यूटिंग क्लस्टर्स सहित विविध बुनियादी ढांचे पर तैनात कर सकते हैं।
Distributed parallel processing framework for numerical computations.
TensorRT एक डीप लर्निंग इन्फरेंस इंजन और सॉफ्टवेयर डेवलपमेंट किट है जिसे NVIDIA GPUs पर उच्च-प्रदर्शन निष्पादन के लिए न्यूरल नेटवर्क को ऑप्टिमाइज़ और डिप्लॉय करने के लिए डिज़ाइन किया गया है। यह एक GPU एक्सेलेरेशन फ्रेमवर्क के रूप में कार्य करता है जो प्रोडक्शन डिप्लॉयमेंट के दौरान प्रशिक्षित मॉडलों के लिए लेटेंसी को कम करता है और थ्रूपुट को बढ़ाता है। यह टूलकिट Open Neural Network Exchange फॉर्मेट से मॉडल इम्पोर्ट करता है और उन्हें ऑप्टिमाइज़्ड इंजनों में बदल देता है। यह ग्राफ-आधारित मॉडल ऑप्टिमाइज़ेशन, लेयर-फ्यूजन कर्नल जनरेशन, और फ्लोटिंग पॉइंट वेट्स को लोअर प्रिसिजन फॉर्मेट में बदलने के लिए प्रिसिजन-आधारित क्वांटिज़ेशन का उपयोग करता है। यह फ्रेमवर्क हार्डवेयर-विशिष्ट इंजन सीरियलाइज़ेशन के लिए क्षमताएं प्रदान करता है और विशेष न्यूरल नेटवर्क लेयर्स के लिए कस्टम प्लगइन्स के माध्यम से इन्फरेंस क्षमताओं के विस्तार का समर्थन करता है।
C++ library for high-performance inference on NVIDIA hardware.
CuPy एक CUDA ऐरे कंप्यूटिंग लाइब्रेरी है जो NVIDIA GPUs पर ऐरे ऑपरेशन्स और संख्यात्मक कंप्यूटिंग को निष्पादित करने के लिए NumPy-संगत इंटरफेस लागू करती है। यह एक GPU-त्वरित संख्यात्मक लाइब्रेरी और CUDA-आधारित SciPy इम्प्लीमेंटेशन के रूप में कार्य करती है, जो वैज्ञानिक और इंजीनियरिंग वर्कलोड के लिए प्रोसेसिंग गति बढ़ाने के लिए ग्राफिक्स हार्डवेयर पर भारी गणनाओं को ऑफलोड करती है। यह लाइब्रेरी मल्टी-फ्रेमवर्क टेंसर एक्सचेंज को सक्षम बनाती है, जिससे मेमोरी कॉपी से बचने के लिए मानकीकृत मेमोरी लेआउट का उपयोग करके विभिन्न डीप लर्निंग फ्रेमवर्क के बीच डेटा बफ़र्स साझा किए जा सकते हैं। यह कस्टम GPU कर्नल एकीकरण का भी समर्थन करती है, जिससे हार्डवेयर निष्पादन पर सटीक नियंत्रण के लिए ऐरे डेटा को लो-लेवल APIs से जोड़ा जा सकता है। व्यापक रूप से, यह प्रोजेक्ट उच्च-प्रदर्शन ऐरे प्रोसेसिंग और वैज्ञानिक कंप्यूटिंग वर्कफ़्लो को कवर करता है। इसकी क्षमताओं में ऐरे कंप्यूटेशन में तेजी लाना और बड़े पैमाने पर संख्यात्मक गणनाओं के लिए उपकरण प्रदान करना शामिल है।
NumPy-compatible multi-dimensional array implementation for CUDA.
Numba एक जस्ट-इन-टाइम कंपाइलर है जो हाई-लेवल Python फंक्शन्स को रनटाइम पर ऑप्टिमाइज़्ड मशीन कोड में अनुवादित करता है। LLVM कंपाइलर इंफ्रास्ट्रक्चर का लाभ उठाकर, यह संख्यात्मक डेटा प्रोसेसिंग और गणितीय गणनाओं में तेजी लाने के लिए एक ढांचा प्रदान करता है, जो स्टेटिकली कंपाइल की गई भाषाओं के बराबर प्रदर्शन स्तर को सक्षम बनाता है। यह प्रोजेक्ट टाइप-इन्फरेंस-आधारित स्पेशलाइजेशन के माध्यम से खुद को अलग करता है, जो निष्पादन के दौरान उपयोग किए जाने वाले विशिष्ट डेटा प्रकारों के अनुरूप मशीन निर्देश उत्पन्न करता है। यह एक लेज़ी कंपाइलेशन पाइपलाइन का उपयोग करता है जो इनवोकेशन के क्षण तक अनुवाद को स्थगित कर देता है, जिससे स्टार्टअप ओवरहेड कम हो जाता है और विविध प्रोसेसर आर्किटेक्चर और ऑपरेटिंग सिस्टम में लगातार प्रदर्शन बना रहता है। कोर कंपाइलेशन के अलावा, यह टूलकिट कई CPU कोर और ग्राफिक्स प्रोसेसिंग यूनिट्स में पुनरावृत्ति संचालन (iterative operations) और ऐरे एक्सप्रेशन्स को वितरित करके हार्डवेयर एक्सेलेरेशन के लिए व्यापक समर्थन प्रदान करता है। यह बड़े पैमाने के संख्यात्मक डेटासेट के लिए थ्रूपुट को अधिकतम करने के लिए वेक्टरइज़ेशन और पैरेललइज़ेशन रणनीतियों का उपयोग करता है, जिससे डेवलपर्स सीधे स्टैंडर्ड कोड से विशेष हार्डवेयर को लक्षित कर सकते हैं।
Compiler for Python array and numerical functions.