3 dépôts
Techniques to increase training speed and reduce memory consumption using hardware-specific acceleration.
Distinct from GPU & Performance: Focuses specifically on improving ML training throughput via mixed precision and hardware acceleration, whereas the parent is a general GPU performance category.
Explore 3 awesome GitHub repositories matching hardware & iot · Training Throughput Optimizations. Refine with filters or upvote what's useful.
This project is a collection of optimized scripts, deployment patterns, and reference implementations designed for scaling and accelerating state-of-the-art AI models. It serves as a multi-domain model zoo and a distributed training framework, providing PyTorch reference implementations for training and deploying models on GPU-accelerated infrastructure. The repository distinguishes itself through an optimization suite focused on NVIDIA GPU hardware, utilizing automatic mixed precision and specialized math modes to increase training speed and throughput. It provides enterprise deployment patt
Implements automatic mixed precision and specialized math modes to increase training speed and throughput on NVIDIA GPUs.
The PyTorch Tutorials repository is a collection of educational resources that provides step-by-step guidance on building, training, and deploying neural networks using the PyTorch framework. It covers the complete machine learning workflow, from data loading and model definition through optimization loops and model persistence, with dedicated guides for distributed training, model fine-tuning, and deployment. The tutorials offer practical demonstrations of adapting pre-trained models to new tasks through transfer learning, scaling training across multiple GPUs or machines using PyTorch's dis
Covers optimizing data loading, memory usage, and gradient flow to maximize training throughput.
sd-scripts is a suite of utilities designed for fine-tuning generative models, preprocessing datasets, and converting model weights. It provides a collection of scripts for executing Stable Diffusion training through methods such as DreamBooth, textual inversion, and full fine-tuning, alongside a framework for creating and managing Low-Rank Adaptation weights. The project features specialized capabilities for model weight conversion between different architectures and precision formats. It includes tools for merging adaptation weights into base models, extracting weights from trained models,
Utilizes specialized precision modes and deep learning primitives to maximize GPU throughput during model training.