2 个仓库
Stores checkpoints in a layout directly loadable by conversion tools for fine-tuning pipelines.
Distinct from Checkpoint Simplifiers: Distinct from Checkpoint Simplifiers: focuses on format compatibility with conversion tools, not size reduction.
Explore 2 awesome GitHub repositories matching artificial intelligence & ml · Compatible Checkpoint Formats. Refine with filters or upvote what's useful.
Torchtitan is a reference implementation for distributed deep learning built within the PyTorch ecosystem. It provides a framework for training large neural network models across multiple GPUs and nodes by combining several parallelism techniques, including fully sharded data parallelism (FSDP), tensor parallelism, and pipeline parallelism, making it possible to train models that exceed the memory capacity of a single device. The system distinguishes itself through asynchronous checkpointing, which saves model and optimizer state to persistent storage without pausing the training loop, enabli
Stores checkpoints in a format directly loadable by downstream conversion tools for fine-tuning.
RLinf is a distributed reinforcement learning orchestrator and embodied AI training framework. It provides the infrastructure to train vision-language-action models and robotic policies using a combination of reinforcement learning and supervised fine-tuning. The system is designed for scaling workloads across GPU clusters, managing the placement of actors, rollout workers, and environment components. It features a specialized robotics data collection pipeline for gathering teleoperated demonstrations and simulation trajectories into standardized replay buffers, alongside a hardware interface
Stores model weights in both HuggingFace and specialized backend formats for interoperability.