awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

2 个仓库

Awesome GitHub RepositoriesDistributed Layer Synchronizers

Systems for coordinating data movement and state consistency between parallelized model layers.

Distinct from Distributed Acceleration Layers: Distinct from Distributed Acceleration Layers: focuses on the synchronization of specific layer data during distributed execution rather than general hardware abstraction.

Explore 2 awesome GitHub repositories matching artificial intelligence & ml · Distributed Layer Synchronizers. Refine with filters or upvote what's useful.

Awesome Distributed Layer Synchronizers GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • zhaochenyang20/awesome-ml-sys-tutorialzhaochenyang20 的头像

    zhaochenyang20/Awesome-ML-SYS-Tutorial

    5,371在 GitHub 上查看↗

    This project provides a comprehensive technical guide and framework for engineering large-scale machine learning systems. It covers the full lifecycle of model development, focusing on the infrastructure and computational principles required to build, train, and serve generative AI models across distributed GPU clusters. The repository distinguishes itself by offering deep-dive tutorials and implementation strategies for complex system challenges. It emphasizes high-performance architectural primitives, such as collective communication orchestration, distributed tensor sharding, and static gr

    Coordinates data movement between independent attention workers and shared MLP layers to maintain state consistency across parallel processing units.

    Python
    在 GitHub 上查看↗5,371
  • b4rtaz/distributed-llamab4rtaz 的头像

    b4rtaz/distributed-llama

    2,837在 GitHub 上查看↗

    Distributed-llama is a distributed inference engine and command line tool for running large language models across multiple networked machines. It functions as a compute cluster manager that coordinates worker nodes to share the computational load of a single model. The system utilizes tensor parallelism to shard model weights across different hosts, allowing the execution of models that exceed the memory capacity of a single piece of hardware. It includes a dedicated format converter to transform standard model files into a compatible binary layout optimized for distributed loading. The eng

    Coordinates the forward pass to ensure each model layer finishes processing across all nodes before the next begins.

    C++distributed-computingdistributed-llmllama2
    在 GitHub 上查看↗2,837
  1. Home
  2. Artificial Intelligence & ML
  3. Distributed Acceleration Layers
  4. Distributed Layer Synchronizers