1 个仓库
Systems that save model and optimizer state asynchronously for interruption‑resilient training in PyTorch.
Distinct from PyTorch Training Frameworks: Distinct from PyTorch Training Frameworks: specifically focuses on asynchronous checkpoint management rather than general training orchestration.
Explore 1 awesome GitHub repository matching artificial intelligence & ml · Asynchronous Checkpoint Managers. Refine with filters or upvote what's useful.
Torchtitan is a reference implementation for distributed deep learning built within the PyTorch ecosystem. It provides a framework for training large neural network models across multiple GPUs and nodes by combining several parallelism techniques, including fully sharded data parallelism (FSDP), tensor parallelism, and pipeline parallelism, making it possible to train models that exceed the memory capacity of a single device. The system distinguishes itself through asynchronous checkpointing, which saves model and optimizer state to persistent storage without pausing the training loop, enabli
Provides asynchronous checkpointing for interruption‑resilient PyTorch training.