awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

86 个仓库

Awesome GitHub RepositoriesModel Deployment Toolkits

Toolkits that streamline the packaging, configuration, and deployment of machine learning models into production environments.

Explore 86 awesome GitHub repositories matching artificial intelligence & ml · Model Deployment Toolkits. Refine with filters or upvote what's useful.

Awesome Model Deployment Toolkits GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • facebookresearch/llamafacebookresearch 的头像

    facebookresearch/llama

    59,466在 GitHub 上查看↗

    Llama is a large language model runtime and inference engine designed to load and execute autoregressive transformer models. It enables the generation of natural language text completions from prompts using pretrained weights. The system features multi-GPU model parallelism, which distributes model weights and workloads across multiple graphics processors to support larger parameter counts. It also incorporates a content safety filter that uses classifiers to intercept and block unsafe inputs or outputs during the inference process. The project covers broad capabilities in distributed model

    Distributes model weights and workloads across multiple graphics processors to handle large parameter counts.

    Python
    在 GitHub 上查看↗59,466
  • ultralytics/ultralyticsultralytics 的头像

    ultralytics/ultralytics

    58,468在 GitHub 上查看↗

    Ultralytics is a comprehensive computer vision framework designed for training, validating, and deploying deep learning models across a wide range of visual recognition tasks. It provides a unified interface for core operations including object detection, instance segmentation, pose estimation, and image classification. By utilizing a modular architecture, the platform allows users to swap model components to balance inference speed and accuracy requirements for diverse applications. The framework distinguishes itself through its support for real-time processing and flexible deployment. It in

    Exports and optimizes models for high-performance execution across cloud and edge hardware environments.

    Pythonclicomputer-visiondeep-learning
    在 GitHub 上查看↗58,468
  • openbmb/voxcpmOpenBMB 的头像

    OpenBMB/VoxCPM

    29,985在 GitHub 上查看↗

    VoxCPM is a multilingual speech synthesis system and text-to-speech inference server. It functions as an AI voice cloning tool and a synthetic voice designer, capable of generating natural speech across global languages and regional dialects using a GPU-accelerated audio generator. The project features a speech model fine-tuning framework that supports both full parameter updates and low-rank adaptation for customizing voice characteristics. It enables high-fidelity voice cloning from reference audio, including cross-lingual voice transfer and acoustic environment mimicry, as well as the crea

    Implements a standardized model format that enables speech synthesis across diverse CPU and GPU backends.

    Pythonaudiodeeplearningminicpm
    在 GitHub 上查看↗29,985
  • meta-llama/llama3meta-llama 的头像

    meta-llama/llama3

    29,254在 GitHub 上查看↗

    Llama 3 is a collection of pretrained, autoregressive transformer-based models designed for natural language generation, reasoning, and complex instruction following. It functions as a generative AI framework that provides the infrastructure for managing model weights, executing neural network inference, and handling computational workloads across diverse knowledge domains. The project distinguishes itself through an integrated AI safety toolkit that employs secondary classification filtering to inspect inputs and outputs, ensuring adherence to usage compliance and safety standards. It suppor

    Supports distributed model deployment by utilizing sharding techniques to split neural network parameters across multiple hardware devices.

    Python
    在 GitHub 上查看↗29,254
  • sgl-project/sglangsgl-project 的头像

    sgl-project/sglang

    29,079在 GitHub 上查看↗

    Sglang is a high-performance inference engine and serving system designed for large language and multimodal models. It provides a programmable interface for orchestrating complex generation workflows, enabling developers to coordinate multi-turn dialogues, tool invocations, and reasoning chains through a domain-specific language. The platform is built to support production-scale deployments, offering an OpenAI-compatible API that allows for integration with existing application ecosystems. The system distinguishes itself through a disaggregated architecture that separates compute-intensive pr

    Separates compute-intensive prompt processing from memory-intensive token generation across distinct hardware nodes.

    Pythonattentionblackwellcuda
    在 GitHub 上查看↗29,079
  • baidu/paddlebaidu 的头像

    baidu/paddle

    23,959在 GitHub 上查看↗

    Paddle is a deep learning framework designed for building, training, and deploying large-scale machine learning models. It incorporates a distributed training engine for optimizing performance across multiple chips and a model inference engine for transforming trained models into production-ready formats for cross-platform execution. The platform features a heterogeneous hardware abstraction and a standardized software stack that allows models to run across diverse hardware architectures through a common interface. It also includes a scientific computing library capable of solving complex dif

    Provides toolkits for transforming trained models into production-ready formats for industrial environments.

    C++
    在 GitHub 上查看↗23,959
  • microsoft/unilmmicrosoft 的头像

    microsoft/unilm

    22,030在 GitHub 上查看↗

    This project is a comprehensive framework and toolkit for developing, optimizing, and deploying transformer-based models across multimodal, document intelligence, and natural language processing tasks. It provides a unified neural architecture that processes text, vision, audio, and document layout data through a shared set of weights, enabling researchers and developers to build foundational models that align cross-modal representations. The platform distinguishes itself through advanced training and inference strategies designed for large-scale deep learning. It incorporates specialized mec

    Provides toolkits for efficient sequence-to-sequence decoding and model compression for production environments.

    Pythonbeitbeit-3bitnet
    在 GitHub 上查看↗22,030
  • qwenlm/qwen-7bQwenLM 的头像

    QwenLM/Qwen-7B

    21,343在 GitHub 上查看↗

    Qwen-7B is a pretrained causal language model designed for natural language generation, text processing, and complex reasoning tasks. It is available as an instruction-tuned model optimized for conversational interactions and a tool-use model capable of executing function calls and interacting with external APIs. The project provides a quantized version of the model to reduce GPU memory usage and supports the development of autonomous agents that can execute code and perform functions to complete complex goals. The system covers a wide range of capabilities including model fine-tuning throug

    Enables model execution across diverse compute environments, including CPUs and multiple GPUs.

    Python
    在 GitHub 上查看↗21,343
  • apache/incubator-mxnetapache 的头像

    apache/incubator-mxnet

    20,812在 GitHub 上查看↗

    Apache MXNet is a deep learning framework and distributed machine learning library designed for training and deploying neural networks across distributed systems, mobile devices, and hardware accelerators. It functions as a cross-platform runtime and a dynamic dataflow scheduler that optimizes neural network execution. The framework provides a multi-language API, enabling the development of machine learning models using Python, R, Julia, Scala, Go, and JavaScript. It supports high-performance model training and the scaling of workloads across multiple GPUs and machines. The system covers cap

    Provides utilities for scaling model inference across multiple hardware devices and nodes using parameter sharding.

    C++
    在 GitHub 上查看↗20,812
  • openai/gpt-ossopenai 的头像

    openai/gpt-oss

    20,191在 GitHub 上查看↗

    gpt-oss is an open-weight large language model and reasoning engine designed for complex reasoning and agentic workflows. It functions as an AI agent framework and model serving API, allowing for local deployment and the hosting of standardized interfaces to expose model completions and internal reasoning processes. The project distinguishes itself as a quantized inference engine, utilizing tensor parallelism and weight quantization to run high-parameter models on limited hardware. It features a reasoning model that employs chain-of-thought processing to solve multi-step logical tasks. The s

    Splits large model weights across multiple GPUs using tensor parallelism to enable high-parameter inference on limited hardware.

    Python
    在 GitHub 上查看↗20,191
  • nari-labs/dianari-labs 的头像

    nari-labs/dia

    19,324在 GitHub 上查看↗

    Dia is a generative AI audio tool and text-to-speech synthesis engine designed for the production-ready deployment of machine learning models. It provides a framework for creating lifelike synthetic speech by conditioning generation on reference audio samples to replicate specific vocal characteristics, emotional tones, and delivery styles. The system distinguishes itself through its ability to perform custom voice cloning and precise control over audio output. Users can adjust generation parameters such as temperature and guidance scale to modify the pacing, creativity, and style of the synt

    Streamlines the management and integration of generative AI models into production environments.

    Pythonaiopen-weighttext-to-speech
    在 GitHub 上查看↗19,324
  • jcjohnson/neural-stylejcjohnson 的头像

    jcjohnson/neural-style

    18,288在 GitHub 上查看↗

    This is a PyTorch implementation of a neural style transfer system. It functions as a convolutional neural network image stylizer and artistic style blender designed to combine the content of one image with the artistic style of another. The system supports blending multiple style sources and adjusting the relative weights between content and style reconstruction. It includes capabilities for preserving the original color palette of the content image and adjusting style scales to determine which artistic patterns are transferred. The pipeline enables high-resolution image processing by distr

    Splits heavy neural network computations across multiple graphics cards for high-resolution image synthesis.

    Lua
    在 GitHub 上查看↗18,288
  • pytorch/visionpytorch 的头像

    pytorch/vision

    17,743在 GitHub 上查看↗

    This project is a comprehensive computer vision library for the PyTorch ecosystem, providing a standardized collection of neural network architectures, datasets, and high-performance transformation utilities. It serves as a foundational framework for building, training, and deploying deep learning models, offering a centralized model registry that allows developers to instantiate architectures with pre-trained weights for tasks such as image classification, object detection, and semantic segmentation. The library distinguishes itself through its modular approach to data and compute management

    Orchestrates the execution of machine learning inference workloads across distributed cloud clusters.

    Pythoncomputer-visionmachine-learning
    在 GitHub 上查看↗17,743
  • kvcache-ai/ktransformerskvcache-ai 的头像

    kvcache-ai/ktransformers

    17,288在 GitHub 上查看↗

    Ktransformers is a comprehensive framework designed for the operation, fine-tuning, and serving of large language models. It functions as a heterogeneous inference engine and quantized execution runtime, enabling the deployment of massive models by distributing computational workloads across both CPU and GPU resources. This architecture allows users to bypass local memory constraints, making it possible to run and train models that exceed the capacity of a single device. The project distinguishes itself through specialized support for sparse architectures, particularly mixture-of-experts mode

    Shards model components across multiple devices to minimize peak memory usage during training and inference.

    Python
    在 GitHub 上查看↗17,288
  • infrasys-ai/aisystemInfrasys-AI 的头像

    Infrasys-AI/AISystem

    17,017在 GitHub 上查看↗

    AISystem is a comprehensive AI full-stack infrastructure project covering the entire pipeline from AI chip architecture to high-level training frameworks. It encompasses the development of AI compiler frameworks, inference engines, and distributed training orchestrators designed to coordinate workloads across a heterogeneous compute stack of CPUs, GPUs, and NPUs. The project focuses on the deep integration of software and hardware, employing software-hardware co-design to align tensor layouts with physical memory structures. It provides specialized capabilities for accelerating Transformer mo

    Converts trained models into optimized formats for specific runtime environments to maximize production resource efficiency.

    Jupyter Notebookaiaiinfraaisys
    在 GitHub 上查看↗17,017
  • thudm/chatglm2-6bTHUDM 的头像

    THUDM/ChatGLM2-6B

    15,565在 GitHub 上查看↗

    ChatGLM2-6B is an open-weight large language model designed for natural language conversations and text generation in both English and Chinese. It functions as a bilingual chat model capable of processing and maintaining coherence across text sequences up to 32K tokens. The model is optimized for local deployment through precision quantization, which reduces memory requirements to allow execution on consumer-grade hardware. It supports distributing model weights across multiple graphics cards to handle parameters that exceed the memory of a single device. The project covers capabilities for

    Splits model parameters across multiple GPUs to execute models that exceed the memory of a single device.

    Python
    在 GitHub 上查看↗15,565
  • zai-org/chatglm2-6bzai-org 的头像

    zai-org/ChatGLM2-6B

    15,564在 GitHub 上查看↗

    ChatGLM2-6B is a bilingual chat large language model designed for natural conversation and text generation in both English and Chinese. It functions as a fine-tunable language model that supports updating weights via specialized scripts to adapt to specific datasets and tasks. The project serves as a quantized inference engine and multi-GPU model orchestrator, enabling the execution of large models on consumer-grade hardware. It is capable of processing long context sequences up to 32K tokens to maintain understanding across extended documents. The system covers capabilities for multilingual

    Splits model parameters across multiple graphics cards to allow large models to fit in available memory.

    Pythonchatglmchatglm-6blarge-language-models
    在 GitHub 上查看↗15,564
  • paddlepaddle/paddledetectionPaddlePaddle 的头像

    PaddlePaddle/PaddleDetection

    14,243在 GitHub 上查看↗

    PaddleDetection is an object detection framework designed for the end-to-end development, training, and deployment of computer vision models. It provides a comprehensive library of modular neural network architectures and pipelines that support object detection, instance segmentation, and multi-object tracking tasks. The project distinguishes itself through a configuration-driven approach that decouples model components like backbones and heads, allowing for the flexible assembly of custom vision workflows. It incorporates advanced techniques such as anchor-free detection logic, joint detecti

    Provides a toolkit for exporting and optimizing neural networks for production inference across diverse hardware.

    Pythonblazefacedeepsortdetr
    在 GitHub 上查看↗14,243
  • zai-org/chatglm3zai-org 的头像

    zai-org/ChatGLM3

    13,764在 GitHub 上查看↗

    ChatGLM3 is a comprehensive framework for deploying, fine-tuning, and serving large language models. It functions as a high-performance inference engine designed to support conversational AI, enabling developers to build interactive agents capable of multi-turn dialogue, autonomous code execution, and structured tool invocation. The project distinguishes itself through its focus on hardware-agnostic deployment and resource optimization. It supports distributed model parallelism across multiple graphics cards, paged key-value caching for concurrent request processing, and weight quantization t

    Enables inference on large models by splitting parameters across multiple graphics cards.

    Python
    在 GitHub 上查看↗13,764
  • thudm/chatglm3THUDM 的头像

    THUDM/ChatGLM3

    13,676在 GitHub 上查看↗

    ChatGLM3 is an open-weights large language model designed for bilingual conversational interactions in English and Chinese. It functions as a tool-augmented system capable of calling external functions and executing internal code to resolve complex tasks. The model utilizes four-bit quantization to reduce memory requirements, enabling inference on consumer hardware and diverse processing units including GPUs and CPUs. It features an expanded context window for processing and summarizing long documents and includes a supervised fine-tuning pipeline for adapting the model to specialized domains

    Executes inference across diverse hardware architectures including GPUs, CPUs, and specialized silicon.

    Python
    在 GitHub 上查看↗13,676
上一个1234…5下一个
  1. Home
  2. Artificial Intelligence & ML
  3. Model Optimization
  4. Inference & Deployment
  5. Model Deployment Toolkits

探索子标签

  • Distributed Deployment Utilities5 个子标签Tools and techniques for scaling model inference across multiple hardware devices using parameter sharding. **Distinct from Model Deployment Toolkits:** Distinct from general deployment toolkits: focuses specifically on distributed sharding and multi-node scaling for large models.
  • Hardware-Agnostic Deployment2 个子标签Strategies for executing models across diverse hardware architectures. **Distinct from Model Deployment Toolkits:** Distinct from general toolkits: focuses on cross-hardware portability and performance optimization.
  • Hardware-Aware DeploymentDeployment systems that automatically detect host hardware capabilities to select and pull the most optimized model image. **Distinct from Hardware-Agnostic Deployment:** Distinct from Hardware-Agnostic Deployment: focuses on active hardware detection and specific image selection rather than generic portability across architectures.