awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

86 repositorios

Awesome GitHub RepositoriesModel Deployment Toolkits

Toolkits that streamline the packaging, configuration, and deployment of machine learning models into production environments.

Explore 86 awesome GitHub repositories matching artificial intelligence & ml · Model Deployment Toolkits. Refine with filters or upvote what's useful.

Awesome Model Deployment Toolkits GitHub Repositories

Encuentra los mejores repositorios con IA.Buscaremos los repositorios que mejor coincidan usando IA.
  • facebookresearch/llamaAvatar de facebookresearch

    facebookresearch/llama

    59,466Ver en GitHub↗

    Llama is a large language model runtime and inference engine designed to load and execute autoregressive transformer models. It enables the generation of natural language text completions from prompts using pretrained weights. The system features multi-GPU model parallelism, which distributes model weights and workloads across multiple graphics processors to support larger parameter counts. It also incorporates a content safety filter that uses classifiers to intercept and block unsafe inputs or outputs during the inference process. The project covers broad capabilities in distributed model

    Distributes model weights and workloads across multiple graphics processors to handle large parameter counts.

    Python
    Ver en GitHub↗59,466
  • ultralytics/ultralyticsAvatar de ultralytics

    ultralytics/ultralytics

    58,468Ver en GitHub↗

    Ultralytics is a comprehensive computer vision framework designed for training, validating, and deploying deep learning models across a wide range of visual recognition tasks. It provides a unified interface for core operations including object detection, instance segmentation, pose estimation, and image classification. By utilizing a modular architecture, the platform allows users to swap model components to balance inference speed and accuracy requirements for diverse applications. The framework distinguishes itself through its support for real-time processing and flexible deployment. It in

    Exports and optimizes models for high-performance execution across cloud and edge hardware environments.

    Pythonclicomputer-visiondeep-learning
    Ver en GitHub↗58,468
  • openbmb/voxcpmAvatar de OpenBMB

    OpenBMB/VoxCPM

    29,985Ver en GitHub↗

    VoxCPM is a multilingual speech synthesis system and text-to-speech inference server. It functions as an AI voice cloning tool and a synthetic voice designer, capable of generating natural speech across global languages and regional dialects using a GPU-accelerated audio generator. The project features a speech model fine-tuning framework that supports both full parameter updates and low-rank adaptation for customizing voice characteristics. It enables high-fidelity voice cloning from reference audio, including cross-lingual voice transfer and acoustic environment mimicry, as well as the crea

    Implements a standardized model format that enables speech synthesis across diverse CPU and GPU backends.

    Pythonaudiodeeplearningminicpm
    Ver en GitHub↗29,985
  • meta-llama/llama3Avatar de meta-llama

    meta-llama/llama3

    29,254Ver en GitHub↗

    Llama 3 is a collection of pretrained, autoregressive transformer-based models designed for natural language generation, reasoning, and complex instruction following. It functions as a generative AI framework that provides the infrastructure for managing model weights, executing neural network inference, and handling computational workloads across diverse knowledge domains. The project distinguishes itself through an integrated AI safety toolkit that employs secondary classification filtering to inspect inputs and outputs, ensuring adherence to usage compliance and safety standards. It suppor

    Supports distributed model deployment by utilizing sharding techniques to split neural network parameters across multiple hardware devices.

    Python
    Ver en GitHub↗29,254
  • sgl-project/sglangAvatar de sgl-project

    sgl-project/sglang

    29,079Ver en GitHub↗

    Sglang is a high-performance inference engine and serving system designed for large language and multimodal models. It provides a programmable interface for orchestrating complex generation workflows, enabling developers to coordinate multi-turn dialogues, tool invocations, and reasoning chains through a domain-specific language. The platform is built to support production-scale deployments, offering an OpenAI-compatible API that allows for integration with existing application ecosystems. The system distinguishes itself through a disaggregated architecture that separates compute-intensive pr

    Separates compute-intensive prompt processing from memory-intensive token generation across distinct hardware nodes.

    Pythonattentionblackwellcuda
    Ver en GitHub↗29,079
  • baidu/paddleAvatar de baidu

    baidu/paddle

    23,959Ver en GitHub↗

    Paddle is a deep learning framework designed for building, training, and deploying large-scale machine learning models. It incorporates a distributed training engine for optimizing performance across multiple chips and a model inference engine for transforming trained models into production-ready formats for cross-platform execution. The platform features a heterogeneous hardware abstraction and a standardized software stack that allows models to run across diverse hardware architectures through a common interface. It also includes a scientific computing library capable of solving complex dif

    Provides toolkits for transforming trained models into production-ready formats for industrial environments.

    C++
    Ver en GitHub↗23,959
  • microsoft/unilmAvatar de microsoft

    microsoft/unilm

    22,030Ver en GitHub↗

    This project is a comprehensive framework and toolkit for developing, optimizing, and deploying transformer-based models across multimodal, document intelligence, and natural language processing tasks. It provides a unified neural architecture that processes text, vision, audio, and document layout data through a shared set of weights, enabling researchers and developers to build foundational models that align cross-modal representations. The platform distinguishes itself through advanced training and inference strategies designed for large-scale deep learning. It incorporates specialized mec

    Provides toolkits for efficient sequence-to-sequence decoding and model compression for production environments.

    Pythonbeitbeit-3bitnet
    Ver en GitHub↗22,030
  • qwenlm/qwen-7bAvatar de QwenLM

    QwenLM/Qwen-7B

    21,343Ver en GitHub↗

    Qwen-7B is a pretrained causal language model designed for natural language generation, text processing, and complex reasoning tasks. It is available as an instruction-tuned model optimized for conversational interactions and a tool-use model capable of executing function calls and interacting with external APIs. The project provides a quantized version of the model to reduce GPU memory usage and supports the development of autonomous agents that can execute code and perform functions to complete complex goals. The system covers a wide range of capabilities including model fine-tuning throug

    Enables model execution across diverse compute environments, including CPUs and multiple GPUs.

    Python
    Ver en GitHub↗21,343
  • apache/incubator-mxnetAvatar de apache

    apache/incubator-mxnet

    20,812Ver en GitHub↗

    Apache MXNet is a deep learning framework and distributed machine learning library designed for training and deploying neural networks across distributed systems, mobile devices, and hardware accelerators. It functions as a cross-platform runtime and a dynamic dataflow scheduler that optimizes neural network execution. The framework provides a multi-language API, enabling the development of machine learning models using Python, R, Julia, Scala, Go, and JavaScript. It supports high-performance model training and the scaling of workloads across multiple GPUs and machines. The system covers cap

    Provides utilities for scaling model inference across multiple hardware devices and nodes using parameter sharding.

    C++
    Ver en GitHub↗20,812
  • openai/gpt-ossAvatar de openai

    openai/gpt-oss

    20,191Ver en GitHub↗

    gpt-oss is an open-weight large language model and reasoning engine designed for complex reasoning and agentic workflows. It functions as an AI agent framework and model serving API, allowing for local deployment and the hosting of standardized interfaces to expose model completions and internal reasoning processes. The project distinguishes itself as a quantized inference engine, utilizing tensor parallelism and weight quantization to run high-parameter models on limited hardware. It features a reasoning model that employs chain-of-thought processing to solve multi-step logical tasks. The s

    Splits large model weights across multiple GPUs using tensor parallelism to enable high-parameter inference on limited hardware.

    Python
    Ver en GitHub↗20,191
  • nari-labs/diaAvatar de nari-labs

    nari-labs/dia

    19,324Ver en GitHub↗

    Dia is a generative AI audio tool and text-to-speech synthesis engine designed for the production-ready deployment of machine learning models. It provides a framework for creating lifelike synthetic speech by conditioning generation on reference audio samples to replicate specific vocal characteristics, emotional tones, and delivery styles. The system distinguishes itself through its ability to perform custom voice cloning and precise control over audio output. Users can adjust generation parameters such as temperature and guidance scale to modify the pacing, creativity, and style of the synt

    Streamlines the management and integration of generative AI models into production environments.

    Pythonaiopen-weighttext-to-speech
    Ver en GitHub↗19,324
  • jcjohnson/neural-styleAvatar de jcjohnson

    jcjohnson/neural-style

    18,288Ver en GitHub↗

    This is a PyTorch implementation of a neural style transfer system. It functions as a convolutional neural network image stylizer and artistic style blender designed to combine the content of one image with the artistic style of another. The system supports blending multiple style sources and adjusting the relative weights between content and style reconstruction. It includes capabilities for preserving the original color palette of the content image and adjusting style scales to determine which artistic patterns are transferred. The pipeline enables high-resolution image processing by distr

    Splits heavy neural network computations across multiple graphics cards for high-resolution image synthesis.

    Lua
    Ver en GitHub↗18,288
  • pytorch/visionAvatar de pytorch

    pytorch/vision

    17,743Ver en GitHub↗

    This project is a comprehensive computer vision library for the PyTorch ecosystem, providing a standardized collection of neural network architectures, datasets, and high-performance transformation utilities. It serves as a foundational framework for building, training, and deploying deep learning models, offering a centralized model registry that allows developers to instantiate architectures with pre-trained weights for tasks such as image classification, object detection, and semantic segmentation. The library distinguishes itself through its modular approach to data and compute management

    Orchestrates the execution of machine learning inference workloads across distributed cloud clusters.

    Pythoncomputer-visionmachine-learning
    Ver en GitHub↗17,743
  • kvcache-ai/ktransformersAvatar de kvcache-ai

    kvcache-ai/ktransformers

    17,288Ver en GitHub↗

    Ktransformers is a comprehensive framework designed for the operation, fine-tuning, and serving of large language models. It functions as a heterogeneous inference engine and quantized execution runtime, enabling the deployment of massive models by distributing computational workloads across both CPU and GPU resources. This architecture allows users to bypass local memory constraints, making it possible to run and train models that exceed the capacity of a single device. The project distinguishes itself through specialized support for sparse architectures, particularly mixture-of-experts mode

    Shards model components across multiple devices to minimize peak memory usage during training and inference.

    Python
    Ver en GitHub↗17,288
  • infrasys-ai/aisystemAvatar de Infrasys-AI

    Infrasys-AI/AISystem

    17,017Ver en GitHub↗

    AISystem is a comprehensive AI full-stack infrastructure project covering the entire pipeline from AI chip architecture to high-level training frameworks. It encompasses the development of AI compiler frameworks, inference engines, and distributed training orchestrators designed to coordinate workloads across a heterogeneous compute stack of CPUs, GPUs, and NPUs. The project focuses on the deep integration of software and hardware, employing software-hardware co-design to align tensor layouts with physical memory structures. It provides specialized capabilities for accelerating Transformer mo

    Converts trained models into optimized formats for specific runtime environments to maximize production resource efficiency.

    Jupyter Notebookaiaiinfraaisys
    Ver en GitHub↗17,017
  • thudm/chatglm2-6bAvatar de THUDM

    THUDM/ChatGLM2-6B

    15,565Ver en GitHub↗

    ChatGLM2-6B is an open-weight large language model designed for natural language conversations and text generation in both English and Chinese. It functions as a bilingual chat model capable of processing and maintaining coherence across text sequences up to 32K tokens. The model is optimized for local deployment through precision quantization, which reduces memory requirements to allow execution on consumer-grade hardware. It supports distributing model weights across multiple graphics cards to handle parameters that exceed the memory of a single device. The project covers capabilities for

    Splits model parameters across multiple GPUs to execute models that exceed the memory of a single device.

    Python
    Ver en GitHub↗15,565
  • zai-org/chatglm2-6bAvatar de zai-org

    zai-org/ChatGLM2-6B

    15,564Ver en GitHub↗

    ChatGLM2-6B is a bilingual chat large language model designed for natural conversation and text generation in both English and Chinese. It functions as a fine-tunable language model that supports updating weights via specialized scripts to adapt to specific datasets and tasks. The project serves as a quantized inference engine and multi-GPU model orchestrator, enabling the execution of large models on consumer-grade hardware. It is capable of processing long context sequences up to 32K tokens to maintain understanding across extended documents. The system covers capabilities for multilingual

    Splits model parameters across multiple graphics cards to allow large models to fit in available memory.

    Pythonchatglmchatglm-6blarge-language-models
    Ver en GitHub↗15,564
  • paddlepaddle/paddledetectionAvatar de PaddlePaddle

    PaddlePaddle/PaddleDetection

    14,243Ver en GitHub↗

    PaddleDetection is an object detection framework designed for the end-to-end development, training, and deployment of computer vision models. It provides a comprehensive library of modular neural network architectures and pipelines that support object detection, instance segmentation, and multi-object tracking tasks. The project distinguishes itself through a configuration-driven approach that decouples model components like backbones and heads, allowing for the flexible assembly of custom vision workflows. It incorporates advanced techniques such as anchor-free detection logic, joint detecti

    Provides a toolkit for exporting and optimizing neural networks for production inference across diverse hardware.

    Pythonblazefacedeepsortdetr
    Ver en GitHub↗14,243
  • zai-org/chatglm3Avatar de zai-org

    zai-org/ChatGLM3

    13,764Ver en GitHub↗

    ChatGLM3 is a comprehensive framework for deploying, fine-tuning, and serving large language models. It functions as a high-performance inference engine designed to support conversational AI, enabling developers to build interactive agents capable of multi-turn dialogue, autonomous code execution, and structured tool invocation. The project distinguishes itself through its focus on hardware-agnostic deployment and resource optimization. It supports distributed model parallelism across multiple graphics cards, paged key-value caching for concurrent request processing, and weight quantization t

    Enables inference on large models by splitting parameters across multiple graphics cards.

    Python
    Ver en GitHub↗13,764
  • thudm/chatglm3Avatar de THUDM

    THUDM/ChatGLM3

    13,676Ver en GitHub↗

    ChatGLM3 is an open-weights large language model designed for bilingual conversational interactions in English and Chinese. It functions as a tool-augmented system capable of calling external functions and executing internal code to resolve complex tasks. The model utilizes four-bit quantization to reduce memory requirements, enabling inference on consumer hardware and diverse processing units including GPUs and CPUs. It features an expanded context window for processing and summarizing long documents and includes a supervised fine-tuning pipeline for adapting the model to specialized domains

    Executes inference across diverse hardware architectures including GPUs, CPUs, and specialized silicon.

    Python
    Ver en GitHub↗13,676
Ant.1234…5Siguiente
  1. Home
  2. Artificial Intelligence & ML
  3. Model Optimization
  4. Inference & Deployment
  5. Model Deployment Toolkits

Explorar subetiquetas

  • Distributed Deployment Utilities5 sub-etiquetasTools and techniques for scaling model inference across multiple hardware devices using parameter sharding. **Distinct from Model Deployment Toolkits:** Distinct from general deployment toolkits: focuses specifically on distributed sharding and multi-node scaling for large models.
  • Hardware-Agnostic Deployment2 sub-etiquetasStrategies for executing models across diverse hardware architectures. **Distinct from Model Deployment Toolkits:** Distinct from general toolkits: focuses on cross-hardware portability and performance optimization.
  • Hardware-Aware DeploymentDeployment systems that automatically detect host hardware capabilities to select and pull the most optimized model image. **Distinct from Hardware-Agnostic Deployment:** Distinct from Hardware-Agnostic Deployment: focuses on active hardware detection and specific image selection rather than generic portability across architectures.