1 repositorio
Analyzes memory usage and communication costs to recommend optimal parallelism configurations.
Distinct from Activation Memory Profilers: Distinct from Activation Memory Profilers: focuses on recommending parallelism configurations based on profile, not just profiling memory.
Explore 1 awesome GitHub repository matching data & databases · Memory-Driven Autotuning. Refine with filters or upvote what's useful.
Torchtitan is a reference implementation for distributed deep learning built within the PyTorch ecosystem. It provides a framework for training large neural network models across multiple GPUs and nodes by combining several parallelism techniques, including fully sharded data parallelism (FSDP), tensor parallelism, and pipeline parallelism, making it possible to train models that exceed the memory capacity of a single device. The system distinguishes itself through asynchronous checkpointing, which saves model and optimizer state to persistent storage without pausing the training loop, enabli
Analyzes memory and communication costs to autotune parallelism configurations.