awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

40 个仓库

Awesome GitHub RepositoriesMulti-GPU Distribution

Techniques for splitting model parameters across multiple graphics cards to overcome memory limitations.

Distinct from Distributed Deployment Utilities: Focuses on the specific capability of multi-GPU sharding for inference, distinct from general distributed deployment utilities.

Explore 40 awesome GitHub repositories matching artificial intelligence & ml · Multi-GPU Distribution. Refine with filters or upvote what's useful.

Awesome Multi-GPU Distribution GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • facebookresearch/llamafacebookresearch 的头像

    facebookresearch/llama

    59,466在 GitHub 上查看↗

    Llama is a large language model runtime and inference engine designed to load and execute autoregressive transformer models. It enables the generation of natural language text completions from prompts using pretrained weights. The system features multi-GPU model parallelism, which distributes model weights and workloads across multiple graphics processors to support larger parameter counts. It also incorporates a content safety filter that uses classifiers to intercept and block unsafe inputs or outputs during the inference process. The project covers broad capabilities in distributed model

    Distributes model weights and workloads across multiple graphics processors to handle large parameter counts.

    Python
    在 GitHub 上查看↗59,466
  • openai/gpt-ossopenai 的头像

    openai/gpt-oss

    20,191在 GitHub 上查看↗

    gpt-oss is an open-weight large language model and reasoning engine designed for complex reasoning and agentic workflows. It functions as an AI agent framework and model serving API, allowing for local deployment and the hosting of standardized interfaces to expose model completions and internal reasoning processes. The project distinguishes itself as a quantized inference engine, utilizing tensor parallelism and weight quantization to run high-parameter models on limited hardware. It features a reasoning model that employs chain-of-thought processing to solve multi-step logical tasks. The s

    Splits large model weights across multiple GPUs using tensor parallelism to enable high-parameter inference on limited hardware.

    Python
    在 GitHub 上查看↗20,191
  • jcjohnson/neural-stylejcjohnson 的头像

    jcjohnson/neural-style

    18,288在 GitHub 上查看↗

    This is a PyTorch implementation of a neural style transfer system. It functions as a convolutional neural network image stylizer and artistic style blender designed to combine the content of one image with the artistic style of another. The system supports blending multiple style sources and adjusting the relative weights between content and style reconstruction. It includes capabilities for preserving the original color palette of the content image and adjusting style scales to determine which artistic patterns are transferred. The pipeline enables high-resolution image processing by distr

    Splits heavy neural network computations across multiple graphics cards for high-resolution image synthesis.

    Lua
    在 GitHub 上查看↗18,288
  • thudm/chatglm2-6bTHUDM 的头像

    THUDM/ChatGLM2-6B

    15,565在 GitHub 上查看↗

    ChatGLM2-6B is an open-weight large language model designed for natural language conversations and text generation in both English and Chinese. It functions as a bilingual chat model capable of processing and maintaining coherence across text sequences up to 32K tokens. The model is optimized for local deployment through precision quantization, which reduces memory requirements to allow execution on consumer-grade hardware. It supports distributing model weights across multiple graphics cards to handle parameters that exceed the memory of a single device. The project covers capabilities for

    Splits model parameters across multiple GPUs to execute models that exceed the memory of a single device.

    Python
    在 GitHub 上查看↗15,565
  • zai-org/chatglm2-6bzai-org 的头像

    zai-org/ChatGLM2-6B

    15,564在 GitHub 上查看↗

    ChatGLM2-6B is a bilingual chat large language model designed for natural conversation and text generation in both English and Chinese. It functions as a fine-tunable language model that supports updating weights via specialized scripts to adapt to specific datasets and tasks. The project serves as a quantized inference engine and multi-GPU model orchestrator, enabling the execution of large models on consumer-grade hardware. It is capable of processing long context sequences up to 32K tokens to maintain understanding across extended documents. The system covers capabilities for multilingual

    Splits model parameters across multiple graphics cards to allow large models to fit in available memory.

    Pythonchatglmchatglm-6blarge-language-models
    在 GitHub 上查看↗15,564
  • zai-org/chatglm3zai-org 的头像

    zai-org/ChatGLM3

    13,764在 GitHub 上查看↗

    ChatGLM3 is a comprehensive framework for deploying, fine-tuning, and serving large language models. It functions as a high-performance inference engine designed to support conversational AI, enabling developers to build interactive agents capable of multi-turn dialogue, autonomous code execution, and structured tool invocation. The project distinguishes itself through its focus on hardware-agnostic deployment and resource optimization. It supports distributed model parallelism across multiple graphics cards, paged key-value caching for concurrent request processing, and weight quantization t

    Enables inference on large models by splitting parameters across multiple graphics cards.

    Python
    在 GitHub 上查看↗13,764
  • thudm/cogvideoTHUDM 的头像

    THUDM/CogVideo

    12,792在 GitHub 上查看↗

    CogVideo is a generative video framework that uses diffusion models and transformer-based architectures to synthesize high-resolution video clips. It functions as both a text-to-video and image-to-video generator, converting textual descriptions or static images into temporal visual sequences. The system integrates large language model capabilities to expand short user prompts into detailed descriptions for better visual alignment. It supports the animation of static images through latent seeding and provides the ability to extend the length of existing video sequences. The project includes

    Distributes model weights across multiple GPUs to enable the generation of high-resolution video.

    Python
    在 GitHub 上查看↗12,792
  • zai-org/cogvideozai-org 的头像

    zai-org/CogVideo

    12,790在 GitHub 上查看↗

    CogVideo is a video generation framework and large language model architecture designed for synthesizing high-resolution video clips from natural language descriptions and images. It functions as a text-to-video and image-to-video generator, while also providing a model for video captioning to analyze visual content into descriptive text summaries. The system supports animating static images into motion sequences and transforming series of images into video based on prompts. It includes capabilities for extending the length of generated video clips to create longer sequences of motion. The f

    Supports splitting model parameters across multiple GPUs to handle large weights and increase throughput during inference.

    Pythoncogvideoximage-to-videollm
    在 GitHub 上查看↗12,790
  • pku-yuangroup/open-sora-planPKU-YuanGroup 的头像

    PKU-YuanGroup/Open-Sora-Plan

    12,163在 GitHub 上查看↗

    Open-Sora-Plan is a text-to-video framework and distributed video training system. It utilizes a diffusion transformer architecture and large language model components to transform written descriptions or image prompts into high-quality video sequences. The system features a distributed infrastructure designed for large-scale video training and inference. It employs sequence parallelism to split high-resolution or long-duration video samples across multiple GPUs and uses a sparse attention mechanism to increase processing speed. The project includes capabilities for both text-to-video and im

    Splits high-resolution video samples across multiple GPUs to accelerate inference through sequence parallelism.

    Python
    在 GitHub 上查看↗12,163
  • mistralai/mistral-srcmistralai 的头像

    mistralai/mistral-src

    10,821在 GitHub 上查看↗

    该项目是一个大语言模型推理库和框架,旨在运行用于文本生成、问题解决和编码辅助的模型。它包括一个用于处理图像和文本组合输入的多模态框架,以及一个基于模型推理执行外部工具的工具调用实现。 该系统具有分布式 GPU 推理引擎,可将大型模型工作负载分散到多个图形处理器上,以提高处理速度并满足内存需求。它还通过预打包的镜像和依赖项提供容器化模型部署,以便在隔离环境中运行推理引擎。 该库涵盖了一系列功能,包括多模态输入分析、函数调用集成,以及用于预测缺失代码段的“中间填充”(fill-in-the-middle)编码。它还支持通过命令行界面进行交互式模型聊天,以维持对话会话。

    Employs techniques to split model parameters across multiple graphics cards to overcome memory limitations and increase speed.

    Jupyter Notebook
    在 GitHub 上查看↗10,821
  • openvinotoolkit/openvinoopenvinotoolkit 的头像

    openvinotoolkit/openvino

    10,414在 GitHub 上查看↗

    OpenVINO is an AI inference engine and model serving platform designed to execute optimized deep learning models across CPUs, GPUs, and NPUs through a unified API. It includes a model optimization toolkit for converting, quantizing, and compressing models from various frameworks, alongside a specialized generative AI runtime for large language models. The project distinguishes itself through a plugin-based hardware acceleration layer that maps neural network operations to vendor-specific drivers. It features advanced execution mechanisms such as continuous batching, speculative decoding, and

    Splits models across multiple GPUs to enable the execution of models that exceed the memory of a single card.

    C++aicomputer-visiondeep-learning
    在 GitHub 上查看↗10,414
  • opengvlab/internvlOpenGVLab 的头像

    OpenGVLab/InternVL

    10,061在 GitHub 上查看↗

    InternVL is a vision-language model framework that fuses a visual encoder with a large language model to translate image features into textual tokens for reasoning. It provides a system for multimodal inference and dialogue, enabling the processing of images and text to answer questions or generate descriptions. The project is distinguished by its high-resolution image processing, which uses dynamic tiling to maintain detail for images up to 4K resolution, and its chain-of-thought visual reasoning for solving complex mathematical and spatial problems. It also supports temporal frame sampling

    Splits model layers across multiple GPUs to execute parameters exceeding single-device memory capacity.

    Pythongptgpt-4ogpt-4v
    在 GitHub 上查看↗10,061
  • lostruins/koboldcppLostRuins 的头像

    LostRuins/koboldcpp

    9,511在 GitHub 上查看↗

    KoboldCPP is a local large language model inference engine and GGUF model runner designed to execute quantized models on personal hardware. It functions as a multimodal AI server and API gateway, providing OpenAI-compatible endpoints that allow third-party clients to interact with locally hosted models. The project distinguishes itself as an AI storytelling backend, featuring dedicated tools for long-form narrative management through persistent memory, world lore tracking, and character state management. It further extends its capabilities as a multimodal server capable of processing text, im

    Partitions model tensors across multiple graphics cards to execute models that exceed a single GPU's memory.

    C++gemmaggmlgguf
    在 GitHub 上查看↗9,511
  • intel/ipex-llmintel 的头像

    intel/ipex-llm

    8,836在 GitHub 上查看↗

    Intel XPU LLM Acceleration Library is a toolkit designed to accelerate large language model inference and finetuning on Intel CPUs, GPUs, and NPUs. It provides a distributed inference engine for scaling models across multiple accelerators, a multimodal model runtime for vision and speech tasks, and a low-bit model quantization tool for converting weights into INT4, FP8, and GGUF formats. The project features a parameter-efficient finetuning framework that enables model adaptation using QLoRA and DPO on Intel hardware. It distinguishes itself by providing specialized optimizations for Intel XP

    Allocates model computation across multiple GPUs to handle models exceeding single-device memory.

    Python
    在 GitHub 上查看↗8,836
  • tiiny-ai/powerinferTiiny-AI 的头像

    Tiiny-AI/PowerInfer

    8,714在 GitHub 上查看↗

    PowerInfer is a high-performance local large language model inference engine and sparse inference framework. It provides a runtime for executing models on consumer-grade hardware, utilizing a GPU acceleration backend to optimize tensor operations for graphics processors. The system distinguishes itself through a sparse inference framework that increases generation speed by skipping computations based on activation sparsity in model weights. It includes a GGUF model converter for transforming weights and metadata into a unified binary format, as well as an OpenAI API compatible server for inte

    Splits tensors across multiple available graphics devices to balance the computational load.

    C++large-language-modelsllamallm
    在 GitHub 上查看↗8,714
  • crazyguitar/pysheeetcrazyguitar 的头像

    crazyguitar/pysheeet

    8,150在 GitHub 上查看↗

    pysheeet 是一个技术参考库,提供了一系列精选的代码片段和实现模式,用于高级 Python 开发、系统集成和高性能计算。它充当实现底层网络编程、原生 C 扩展以及异步和并发编程的综合指南。 该项目为大语言模型的开发和部署提供了专门的框架,包括用于分布式 GPU 推理和高性能服务的工具。它还包括用于高性能计算集群编排的详细模式,涵盖 GPU 资源分配和多节点工作负载管理。 该库涵盖了广泛的功能,包括安全网络通信和加密、对象关系映射和数据库管理,以及复杂数据结构和算法的实现。它还提供用于内存管理、通过外部函数接口(FFI)进行原生互操作以及系统级 OS 集成的实用程序。

    Implements strategies for splitting model weights across multiple GPUs using tensor parallelism for high-throughput inference.

    Python
    在 GitHub 上查看↗8,150
  • thudm/cogvlmTHUDM 的头像

    THUDM/CogVLM

    6,742在 GitHub 上查看↗

    CogVLM is a multimodal large language model designed to integrate visual and textual data for reasoning about images and generating natural language. It functions as a visual question answering system that analyzes image content to provide detailed descriptions or answer specific questions. The project includes a visual grounding model capable of mapping text descriptions to precise bounding box coordinates within an image. It also features a vision-based automation agent that analyzes screen captures to generate execution plans and interaction coordinates for software interfaces. The system

    Distributes large model parameters across multiple GPUs to overcome memory limits and reduce latency.

    Python
    在 GitHub 上查看↗6,742
  • elder-plinius/obliteratuselder-plinius 的头像

    elder-plinius/OBLITERATUS

    6,736在 GitHub 上查看↗

    Obliteratus is a weight ablation framework and refusal removal tool designed to identify and delete the internal representations responsible for content refusals in large language models without retraining. It functions as a circuit analysis suite that maps the geometric structure of model guardrails to isolate the specific layers and attention heads that enforce refusals. The project enables the removal of these behaviors through geometric projection, rank-1 adapter ablation for reversible modifications, and the application of steering vectors to alter behavior during inference. It includes

    Implements multi-GPU sharding to overcome memory limitations when removing weights from large models.

    Python
    在 GitHub 上查看↗6,736
  • meta-pytorch/gpt-fastmeta-pytorch 的头像

    meta-pytorch/gpt-fast

    6,223在 GitHub 上查看↗

    gpt-fast 是一个 PyTorch Transformer 推理引擎,专为使用原生张量库实现的文本生成而设计。它提供了一个运行时环境,用于执行大语言模型,而无需外部 C++ 扩展。 该项目实现了推测解码,通过使用小型草稿模型进行 Token 预测和较大模型进行验证来加速生成。它通过编译后的预填充阶段和跨多个图形处理单元分片线性层的多 GPU 张量并行库进一步优化了性能。 内存效率通过支持 int8 和 int4 权重以及分组张量量化的量化运行时进行管理。该系统还包括用于架构参数化、文本分词和使用标准化工具进行模型准确性评估的工具。

    Provides a toolkit for splitting model weights across multiple GPUs using tensor parallelism.

    Python
    在 GitHub 上查看↗6,223
  • nvidia/warpNVIDIA 的头像

    NVIDIA/warp

    6,233在 GitHub 上查看↗

    Warp is a Python framework that JIT-compiles Python functions into CUDA kernels for GPU-accelerated parallel computation, with built-in automatic differentiation and multi-framework array interoperability. At its core, it provides a GPU kernel compilation system that enables writing and executing custom GPU kernels directly from Python, while supporting automatic gradient computation through those kernels for integration with machine learning pipelines. The framework also includes tile-based cooperative computing, where thread blocks partition into tiles for shared-memory and tensor-core opera

    Uses JAX's shard_map to run Warp kernels on sharded arrays across multiple GPUs.

    Pythoncudadifferentiable-programminggpu
    在 GitHub 上查看↗6,233
上一个12下一个
  1. Home
  2. Artificial Intelligence & ML
  3. Model Optimization
  4. Inference & Deployment
  5. Model Deployment Toolkits
  6. Distributed Deployment Utilities
  7. Multi-GPU Distribution

探索子标签

  • FFT DistributionsDistributes Fast Fourier Transform calculations across multiple GPUs in a single node for larger datasets and higher throughput. **Distinct from Multi-GPU Distribution:** Distinct from Multi-GPU Distribution: focuses specifically on distributing FFT computations, not general model or workload distribution.
  • Gaussian Splatting Multi-GPU DistributionsSplitting 3D Gaussian rasterization across multiple GPUs to handle larger scenes and increase throughput. **Distinct from Multi-GPU Distribution:** Distinct from Multi-GPU Distribution: specifically distributes Gaussian splatting rasterization, not general model parameters.
  • Tensor-Parallel Inference Distributions2 个子标签Splitting model weights across multiple GPUs using tensor parallelism to handle models that exceed single-device memory. **Distinct from Multi-GPU Distribution:** Distinct from Multi-GPU Distribution: focuses on tensor parallelism specifically, not general model sharding or data parallelism.