awesome-repositories.com
المدونة
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعحولكيفية ترتيب النتائجالصحافةخادم MCP
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

71 مستودعات

Awesome GitHub RepositoriesModel Serving & Deployment

Tools for deploying, serving, and optimizing AI models in production.

Explore 71 awesome GitHub repositories matching part of an awesome list · Model Serving & Deployment. Refine with filters or upvote what's useful.

Awesome Model Serving & Deployment GitHub Repositories

اعثر على أفضل المستودعات باستخدام الذكاء الاصطناعي.سنبحث عن أفضل المستودعات المطابقة باستخدام الذكاء الاصطناعي.
  • open-webui/open-webuiالصورة الرمزية لـ open-webui

    open-webui/open-webui

    142,694عرض على GitHub↗

    Open WebUI is a self-hosted, web-based platform designed for interacting with local and remote artificial intelligence models. It functions as a unified interface and orchestration suite, enabling users to build, deploy, and manage specialized AI agents equipped with custom instructions, external tool access, and private knowledge bases. The platform distinguishes itself through a modular architecture that supports complex AI workflows. It features a plugin-based framework for custom logic and pipeline-based request processing, allowing developers to filter or transform data streams before th

    Provides a self-hosted, offline-capable AI platform.

    Pythonaillmllm-ui
    عرض على GitHub↗142,694
  • ggml-org/llama.cppالصورة الرمزية لـ ggml-org

    ggml-org/llama.cpp

    116,799عرض على GitHub↗

    Llama.cpp is an inference engine designed for the local execution of text-based and multimodal language models on consumer hardware. It provides a core environment for running models that process both text and image inputs, utilizing hardware-accelerated backends to optimize performance across diverse CPU and GPU architectures. The project distinguishes itself by offering a lightweight HTTP server that adheres to standard API specifications, enabling chat completion, embeddings, and reranking services. It includes a suite of tools for model quantization and conversion, which reduces memory us

    Performs efficient local inference for various LLMs.

    C++ggml
    عرض على GitHub↗116,799
  • vllm-project/vllmالصورة الرمزية لـ vllm-project

    vllm-project/vllm

    83,048عرض على GitHub↗

    vLLM is a high-throughput inference engine designed for the efficient serving and execution of large language models. It functions as a production-ready distributed model server, providing standard API protocols for online serving while also supporting offline batch processing. The system is built to maximize token generation speed and memory efficiency, enabling both large-scale cloud deployments and local execution on personal hardware. The project distinguishes itself through advanced memory management and request scheduling techniques, most notably its use of non-contiguous key-value cach

    Provides a high-throughput, memory-efficient LLM serving engine.

    Pythonamdblackwellcuda
    عرض على GitHub↗83,048
  • berriai/litellmالصورة الرمزية لـ BerriAI

    BerriAI/litellm

    50,579عرض على GitHub↗

    LiteLLM is a unified gateway and proxy server designed to centralize access to over one hundred language model providers. It provides a standardized API interface that abstracts vendor-specific schemas, allowing developers to interact with diverse models through a single, consistent format. By acting as a central traffic management layer, it enables organizations to route, secure, and govern model interactions across multiple deployments. The platform distinguishes itself through its policy-driven architecture, which uses configuration-based routing to manage traffic distribution, load balanc

    Acts as a proxy server for calling multiple LLM APIs.

    Pythonai-gatewayanthropicazure-openai
    عرض على GitHub↗50,579
  • mudler/localaiالصورة الرمزية لـ mudler

    mudler/LocalAI

    46,889عرض على GitHub↗

    LocalAI is a self-hosted inference server that enables the execution of machine learning models directly on local hardware. By providing a unified interface for text, image, and audio processing, it allows users to maintain full control over data privacy and infrastructure costs while eliminating dependencies on external network services. The platform functions as an API gateway that mimics standard cloud-based artificial intelligence interfaces, allowing existing applications to integrate local models as drop-in replacements. It utilizes a container-based architecture to package runtimes and

    Provides an OpenAI-compatible API for local inference.

    Goaiapiaudio-generation
    عرض على GitHub↗46,889
  • exo-explore/exoالصورة الرمزية لـ exo-explore

    exo-explore/exo

    45,380عرض على GitHub↗

    Exo is a distributed inference engine designed to run machine learning models across local hardware. It functions as a network orchestration layer that automatically discovers available devices to form a unified computing cluster, allowing users to scale artificial intelligence workloads by distributing computational tasks across multiple machines. The platform distinguishes itself through its ability to manage the entire lifecycle of local models while providing a standardized gateway for external applications. By translating local model outputs into industry-standard formats, it enables exi

    Runs AI clusters on local consumer hardware.

    Python
    عرض على GitHub↗45,380
  • mindsdb/mindsdbالصورة الرمزية لـ mindsdb

    mindsdb/mindsdb

    39,313عرض على GitHub↗

    MindsDB is an AI-native database engine that treats machine learning models and autonomous agents as virtual tables. By mapping external data sources, predictive models, and third-party services directly into the database schema, it enables users to perform inference, data retrieval, and complex orchestration using standard SQL syntax. The platform distinguishes itself through an autonomous agent orchestrator that executes iterative reasoning loops, allowing agents to plan data access and synthesize natural language responses from connected knowledge bases. It functions as a federated data ga

    Serves and fine-tunes models directly from databases.

    Makefileagentsaianalytics
    عرض على GitHub↗39,313
  • sgl-project/sglangالصورة الرمزية لـ sgl-project

    sgl-project/sglang

    29,079عرض على GitHub↗

    Sglang is a high-performance inference engine and serving system designed for large language and multimodal models. It provides a programmable interface for orchestrating complex generation workflows, enabling developers to coordinate multi-turn dialogues, tool invocations, and reasoning chains through a domain-specific language. The platform is built to support production-scale deployments, offering an OpenAI-compatible API that allows for integration with existing application ecosystems. The system distinguishes itself through a disaggregated architecture that separates compute-intensive pr

    Provides a fast serving framework for LLMs and VLMs.

    Pythonattentionblackwellcuda
    عرض على GitHub↗29,079
  • jina-ai/serveالصورة الرمزية لـ jina-ai

    jina-ai/serve

    21,859عرض على GitHub↗

    Serve is a multimodal AI orchestrator and inference server designed for deploying and scaling machine learning models as cloud-native services. It functions as a containerized workflow engine and distributed service mesh that routes multimodal data through connected execution units. The framework provides specialized capabilities for large language models, including a token streaming gateway that delivers generated text incrementally to reduce perceived latency. It distinguishes itself by enabling the chaining of executors into complex data processing pipelines and the orchestration of these

    Builds AI services with gRPC and HTTP support.

    Pythoncloud-nativecncfdeep-learning
    عرض على GitHub↗21,859
  • vercel/aiالصورة الرمزية لـ vercel

    vercel/ai

    21,885عرض على GitHub↗

    This project is a comprehensive framework for building AI-powered applications, providing a unified toolkit for orchestrating language models, autonomous agents, and interactive user interfaces. It serves as a central library for managing the entire lifecycle of AI interactions, from initial prompt generation and model provider abstraction to complex, multi-step reasoning and tool execution. The framework distinguishes itself through its deep integration with frontend development, specifically by enabling generative user interfaces that render dynamic components directly from model outputs. I

    Builds AI-powered applications using modern web frameworks.

    TypeScriptanthropicartificial-intelligencegemini
    عرض على GitHub↗21,885
  • kvcache-ai/ktransformersالصورة الرمزية لـ kvcache-ai

    kvcache-ai/ktransformers

    17,288عرض على GitHub↗

    Ktransformers is a comprehensive framework designed for the operation, fine-tuning, and serving of large language models. It functions as a heterogeneous inference engine and quantized execution runtime, enabling the deployment of massive models by distributing computational workloads across both CPU and GPU resources. This architecture allows users to bypass local memory constraints, making it possible to run and train models that exceed the capacity of a single device. The project distinguishes itself through specialized support for sparse architectures, particularly mixture-of-experts mode

    Optimizes LLM inference with flexible framework support.

    Python
    عرض على GitHub↗17,288
  • bentoml/openllmالصورة الرمزية لـ bentoml

    bentoml/OpenLLM

    12,115عرض على GitHub↗

    OpenLLM is a framework for deploying, managing, and scaling open-source large language models

    Runs open-source LLMs as OpenAI-compatible APIs.

    Pythonbentomlfine-tuningllama
    عرض على GitHub↗12,115
  • geeeekexplorer/nano-vllmالصورة الرمزية لـ GeeeekExplorer

    GeeeekExplorer/nano-vllm

    11,745عرض على GitHub↗

    Nano-vllm is a high-performance inference engine designed for executing large language models locally. It functions as a specialized runtime that prioritizes accelerated token generation and efficient hardware utilization for text generation tasks. The project distinguishes itself through a comprehensive suite of optimization techniques, including a graph compilation engine that transforms neural network operations into pre-compiled execution plans. It also incorporates a tensor parallelism framework to distribute model weights across multiple hardware accelerators, effectively reducing memor

    Implements a lightweight, fast inference engine.

    Pythondeep-learninginferencellm
    عرض على GitHub↗11,745
  • dataelement/bishengالصورة الرمزية لـ dataelement

    dataelement/bisheng

    11,455عرض على GitHub↗

    Bisheng is an enterprise AI framework and LLM DevOps platform designed to manage the full lifecycle of large language models. It provides a unified system for dataset curation, supervised fine-tuning, model versioning, and performance evaluation. The platform features a visual workflow orchestrator for building retrieval-augmented generation pipelines and complex task sequences using flowcharts with conditional logic and human intervention points. It also includes an AI agent framework that uses a specialized guidance language to embed domain expertise and professional business logic into aut

    Focuses on enterprise-grade LLM application development.

    TypeScript
    عرض على GitHub↗11,455
  • lyogavin/airllmالصورة الرمزية لـ lyogavin

    lyogavin/airllm

    11,508عرض على GitHub↗

    Airllm is a framework designed to execute and fine-tune large language models on consumer-grade hardware. By employing layer-wise model decomposition and memory-efficient loading techniques, the engine enables the operation of massive models that would otherwise exceed available system or video memory. The project distinguishes itself through a suite of optimization strategies that balance memory footprint with performance. It utilizes block-wise weight quantization and asynchronous layer prefetching to reduce resource consumption and hide data transfer latency. Additionally, the framework su

    Optimizes memory usage for running large models on limited hardware.

    Jupyter Notebookchinese-llmchinese-nlpfinetune
    عرض على GitHub↗11,508
  • huggingface/text-generation-inferenceالصورة الرمزية لـ huggingface

    huggingface/text-generation-inference

    10,775عرض على GitHub↗

    Text Generation Inference is a production-ready engine designed for the deployment and serving of large language models. It functions as a containerized runtime environment that manages model execution, scales across distributed hardware, and provides high-performance inference capabilities for demanding production environments. The project distinguishes itself through advanced optimization techniques, including continuous batching to maximize hardware utilization and tensor parallelism to shard large models across multiple accelerator cards. It supports efficient inference through custom com

    Generates text using large language models.

    Pythonbloomdeep-learningfalcon
    عرض على GitHub↗10,775
  • triton-inference-server/serverالصورة الرمزية لـ triton-inference-server

    triton-inference-server/server

    10,768عرض على GitHub↗

    Triton Inference Server is a high-performance server designed to deploy machine learning models from multiple frameworks across GPUs and CPUs. It functions as a hardware-accelerated inference engine and a gRPC inference gateway, providing a standardized communication layer for transmitting binary tensor data with low latency. The system acts as a multi-framework model orchestrator, allowing users to link multiple AI models into ensembles and scripts to create complex inference pipelines. It also serves as a model lifecycle manager, providing controls to load, unload, and monitor the performan

    Maximizes GPU/CPU utilization for model deployment.

    Pythonclouddatacenterdeep-learning
    عرض على GitHub↗10,768
  • openvinotoolkit/openvinoالصورة الرمزية لـ openvinotoolkit

    openvinotoolkit/openvino

    10,414عرض على GitHub↗

    OpenVINO is an AI inference engine and model serving platform designed to execute optimized deep learning models across CPUs, GPUs, and NPUs through a unified API. It includes a model optimization toolkit for converting, quantizing, and compressing models from various frameworks, alongside a specialized generative AI runtime for large language models. The project distinguishes itself through a plugin-based hardware acceleration layer that maps neural network operations to vendor-specific drivers. It features advanced execution mechanisms such as continuous batching, speculative decoding, and

    Downloads, configures, and serves generative AI models directly from the Hugging Face hub without manual preparation.

    C++aicomputer-visiondeep-learning
    عرض على GitHub↗10,414
  • skypilot-org/skypilotالصورة الرمزية لـ skypilot-org

    skypilot-org/skypilot

    10,172عرض على GitHub↗

    SkyPilot is a multi-cloud AI orchestrator and distributed task scheduler designed to launch and manage AI workloads across various cloud providers, Kubernetes, and Slurm clusters. It functions as an infrastructure-as-code framework that uses declarative files to define resource requirements and setup commands for consistent execution across different environments. The project differentiates itself through automated cost optimization, selecting the most affordable GPU or TPU hardware and managing spot instances to reduce expenses. It also provides a remote development environment that bridges

    Runs AI and batch jobs across any cloud provider.

    Python
    عرض على GitHub↗10,172
  • autogluon/autogluonالصورة الرمزية لـ autogluon

    autogluon/autogluon

    9,997عرض على GitHub↗

    AutoGluon is an automated machine learning framework and multimodal library designed to automate the end-to-end pipeline from data preprocessing to high-accuracy model training and validation. It functions as an automated model trainer for tabular, image, text, and time series data, as well as a tool for time series forecasting and foundation model finetuning. The project is distinguished by its ability to jointly process and fuse different data types, allowing for the construction of multimodal neural networks that integrate images, text, and structured tables. It supports zero-shot inferenc

    Packages trained predictors into containers or serverless functions for real-time or batch inference.

    Pythonautogluonautomated-machine-learningautoml
    عرض على GitHub↗9,997
السابق123…4التالي
  1. Home
  2. Part of an Awesome List
  3. AI & Machine Learning
  4. Model Serving & Deployment

استكشف الوسوم الفرعية

  • Cached Model DeploymentsDeployments that leverage pre-downloaded model artifacts stored on local cluster nodes to reduce startup latency. **Distinct from Model Serving & Deployment:** Distinct from general Model Serving & Deployment: focuses on deployments that use locally cached models rather than downloading from remote storage each time.
  • Hub-Integrated Deployment1 وسم فرعيAutomated deployment workflows that download and configure models directly from AI model hubs. **Distinct from Model Serving & Deployment:** Distinct from general Model Serving & Deployment: specifically covers the automated pipeline from a hub like Hugging Face to a serving environment.
  • Instance Configuration StrategiesChooses between isolated single-model instances, shared multi-model instances, or sharded serving for models that exceed single-device memory. **Distinct from Model Serving & Deployment:** Distinct from Model Serving & Deployment: focuses on the specific configuration of model instances, not general deployment.
  • Low-Latency Serving TechniquesDeploys trained models with dynamic batching, memory management, and quantization to achieve high-throughput, low-latency responses. **Distinct from Model Serving & Deployment:** Distinct from Model Serving & Deployment: focuses specifically on low-latency serving techniques, not general deployment.
  • Production Traffic ScalingTechniques for handling high-volume production traffic through batching and concurrent execution. **Distinct from Model Serving & Deployment:** Focuses specifically on request-time scaling techniques like dynamic batching, rather than general deployment tools.
  • Small Model ServingStrategies and instructions for publishing lightweight AI models to serving platforms or application codebases. **Distinct from Model Serving & Deployment:** Specifically targets the deployment of small-parameter models and their associated executable flows, rather than general production serving.