awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

2 个仓库

Awesome GitHub RepositoriesCompatible Checkpoint Formats

Stores checkpoints in a layout directly loadable by conversion tools for fine-tuning pipelines.

Distinct from Checkpoint Simplifiers: Distinct from Checkpoint Simplifiers: focuses on format compatibility with conversion tools, not size reduction.

Explore 2 awesome GitHub repositories matching artificial intelligence & ml · Compatible Checkpoint Formats. Refine with filters or upvote what's useful.

Awesome Compatible Checkpoint Formats GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • pytorch/torchtitanpytorch 的头像

    pytorch/torchtitan

    5,084在 GitHub 上查看↗

    Torchtitan is a reference implementation for distributed deep learning built within the PyTorch ecosystem. It provides a framework for training large neural network models across multiple GPUs and nodes by combining several parallelism techniques, including fully sharded data parallelism (FSDP), tensor parallelism, and pipeline parallelism, making it possible to train models that exceed the memory capacity of a single device. The system distinguishes itself through asynchronous checkpointing, which saves model and optimizer state to persistent storage without pausing the training loop, enabli

    Stores checkpoints in a format directly loadable by downstream conversion tools for fine-tuning.

    Python
    在 GitHub 上查看↗5,084
  • rlinf/rlinfRLinf 的头像

    RLinf/RLinf

    2,502在 GitHub 上查看↗

    RLinf is a distributed reinforcement learning orchestrator and embodied AI training framework. It provides the infrastructure to train vision-language-action models and robotic policies using a combination of reinforcement learning and supervised fine-tuning. The system is designed for scaling workloads across GPU clusters, managing the placement of actors, rollout workers, and environment components. It features a specialized robotics data collection pipeline for gathering teleoperated demonstrations and simulation trajectories into standardized replay buffers, alongside a hardware interface

    Stores model weights in both HuggingFace and specialized backend formats for interoperability.

    Pythonagentic-aiembodied-aireinforcement-learning
    在 GitHub 上查看↗2,502
  1. Home
  2. Artificial Intelligence & ML
  3. Model Checkpointing
  4. Compatible Checkpoint Formats