awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

400 रिपॉजिटरी

Awesome GitHub RepositoriesModel Inference and Serving

Platforms and techniques for deploying, optimizing, and serving machine learning models for production use.

Explore 400 awesome GitHub repositories matching artificial intelligence & ml · Model Inference and Serving. Refine with filters or upvote what's useful.

Awesome Model Inference and Serving GitHub Repositories

AI के साथ बेहतरीन रिपॉजिटरी खोजें।हम AI का उपयोग करके सबसे सटीक रिपॉजिटरी खोजेंगे।
  • tensorflow/tensorflowtensorflow का अवतार

    tensorflow/tensorflow

    195,697GitHub पर देखें↗

    TensorFlow is a comprehensive machine learning framework designed for the construction, training, and deployment of complex mathematical models. It utilizes a graph-based execution model that represents operations as directed acyclic graphs, enabling automatic differentiation and efficient parallel processing. The system provides high-level interfaces for defining neural network architectures, alongside a robust engine for managing multidimensional array structures and tensor mathematics. The framework distinguishes itself through a scalable distributed runtime that orchestrates workloads acr

    Optimizes execution performance by setting specific model weights to zero through target-aware authoring and specialized kernels.

    C++deep-learningdeep-neural-networksdistributed
    GitHub पर देखें↗195,697
  • jmorganca/ollamajmorganca का अवतार

    jmorganca/ollama

    174,350GitHub पर देखें↗

    Ollama is a cross-platform runtime for managing, serving, and executing large language models on local hardware. It functions as a model manager and orchestrator that allows for the downloading, updating, and organization of model weights and configurations to ensure private and offline inference. The system provides a local inference API and a RESTful interface for programmatic model lifecycle management and text generation. It utilizes a compiled C++ backend to handle tensor operations and memory management. To support various hardware configurations, the runtime employs dynamic GPU offloa

    Facilitates the downloading and execution of language models on local computing environments for private inference.

    Go
    GitHub पर देखें↗174,350
  • huggingface/pytorch-pretrained-berthuggingface का अवतार

    huggingface/pytorch-pretrained-BERT

    161,658GitHub पर देखें↗

    This project is a PyTorch transformer model library and pre-trained model framework. It serves as a deep learning model hub and multimodal inference engine, providing a centralized system for loading, executing, and fine-tuning state-of-the-art model checkpoints. The library focuses on multimodal machine learning, enabling predictions across text, vision, and audio data. It provides specialized capabilities for model framework interoperability, allowing the conversion of weights and definitions between different deep learning libraries. The platform covers the full model lifecycle, including

    Provides a framework for loading models and generating predictions across text, vision, audio, and multimodal data.

    Python
    GitHub पर देखें↗161,658
  • huggingface/transformershuggingface का अवतार

    huggingface/transformers

    161,630GitHub पर देखें↗

    Transformers is a comprehensive library for machine learning that provides a unified interface for training, fine-tuning, and deploying transformer-based models. It supports a wide range of tasks, including text classification, language modeling, question answering, and sequence-to-sequence translation, while offering specialized architectures for both text and vision processing. The framework includes tools for managing the entire model lifecycle, from data preprocessing and tokenization to distributed training and inference. The library features extensive support for model optimization and

    Extends standard library functionality with specialized loaders for device mapping, quantization, and custom attention backends.

    Pythonaudiodeep-learningdeepseek
    GitHub पर देखें↗161,630
  • ggerganov/llama.cppggerganov का अवतार

    ggerganov/llama.cpp

    116,912GitHub पर देखें↗

    llama.cpp is a high-performance C++ inference engine and runtime for executing large language models locally across various hardware architectures. It provides the core components for local model execution, including a dedicated model quantizer for compressing weights into the GGUF format and a system for generating text embeddings for semantic search. The project distinguishes itself through specialized memory and execution optimizations, such as block-wise weight quantization to reduce memory footprints and memory-mapped model loading. It supports structured text generation by using formal

    Provides a high-performance runtime for deploying and executing models across diverse local hardware architectures.

    C++
    GitHub पर देखें↗116,912
  • immich-app/immichimmich-app का अवतार

    immich-app/immich

    104,236GitHub पर देखें↗

    Immich is a self-hosted media management platform designed to provide a centralized, private repository for photos and videos. It functions as a comprehensive system for organizing, backing up, and viewing personal media collections across mobile devices, web browsers, and external storage locations. By maintaining full control over data ownership and storage infrastructure, the platform ensures that users retain sovereignty over their digital assets. The system distinguishes itself through a distributed architecture that coordinates background media synchronization, real-time filesystem moni

    Processes machine learning tasks using externalized models and thread pools to optimize performance for image and text analysis.

    TypeScriptbackup-toolfluttergoogle-photos
    GitHub पर देखें↗104,236
  • browser-use/browser-usebrowser-use का अवतार

    browser-use/browser-use

    100,229GitHub पर देखें↗

    Browser-use is a framework for building autonomous agents that navigate, interact with, and extract data from web interfaces using natural language instructions. By acting as an orchestration layer between large language models and browser automation protocols, it enables the execution of complex, multi-step workflows without relying on brittle selectors. The system functions as a headless browser controller, providing a programmatic interface to manage browser instances and execute granular interactions. The project distinguishes itself through its ability to translate high-level intent into

    Manages settings and parameters for integrating specific generative AI models into browser-based automation workflows.

    Pythonai-agentsai-toolsbrowser-automation
    GitHub पर देखें↗100,229
  • hacksider/deep-live-camhacksider का अवतार

    hacksider/Deep-Live-Cam

    93,878GitHub पर देखें↗

    Deep-Live-Cam is a generative video transformation tool designed for real-time facial manipulation and cinematic enhancement. It functions as a local-first AI runtime, performing all media processing directly on the user's hardware to ensure complete data privacy without external network dependencies. By utilizing a high-performance processing pipeline, the application enables live face swapping and interactive video modifications during active streaming sessions or on pre-recorded media. The system distinguishes itself through a hardware-abstraction execution layer that dynamically routes co

    Executes deep learning models directly on hardware-specific providers to minimize latency.

    Pythonaiai-deep-fakeai-face
    GitHub पर देखें↗93,878
  • punkpeye/awesome-mcp-serverspunkpeye का अवतार

    punkpeye/awesome-mcp-servers

    89,264GitHub पर देखें↗

    This project serves as a centralized directory and interoperability hub for the Model Context Protocol, providing a curated collection of standardized service connectors that bridge artificial intelligence models with external software, databases, and APIs. It facilitates the integration of AI agents with diverse ecosystems by offering a registry of machine-readable interface definitions that enable dynamic tool discovery and structured context injection. The directory distinguishes itself by focusing on the protocol-based interoperability required for autonomous AI agents to interact with he

    Delivers standardized interfaces for agents to control desktop environments, manage windows, and simulate user input.

    aimcp
    GitHub पर देखें↗89,264
  • opencv/opencvopencv का अवतार

    opencv/opencv

    89,201GitHub पर देखें↗

    OpenCV is a comprehensive computer vision library designed for real-time performance and cross-platform deployment. It provides a native execution environment that leverages multi-threaded operations and automated memory management to handle intensive computational tasks, including image processing and machine learning model inference. The library distinguishes itself through a data-oriented matrix framework that utilizes proxy-based array abstractions to provide a consistent interface for multidimensional data. By employing factory-pattern algorithm interfaces and runtime type dispatching, i

    Executes pre-trained neural networks to perform classification, detection, and segmentation tasks on visual data.

    C++c-plus-pluscomputer-visiondeep-learning
    GitHub पर देखें↗89,201
  • modelcontextprotocol/serversmodelcontextprotocol का अवतार

    modelcontextprotocol/servers

    87,320GitHub पर देखें↗

    The Model Context Protocol is a standardized communication framework designed to connect language models to external data sources, functional tools, and interactive user interfaces. It provides a vendor-neutral interface layer that enables AI hosts to discover and execute capabilities across heterogeneous service environments, using a JSON-RPC based messaging standard to facilitate bidirectional communication between clients and servers. The protocol distinguishes itself through a robust capability-based handshake that negotiates feature sets during session initialization, ensuring compatibil

    Creates a unified interface layer that enables seamless interaction between diverse AI clients and backend service providers.

    TypeScript
    GitHub पर देखें↗87,320
  • zed-industries/zedzed-industries का अवतार

    zed-industries/zed

    85,338GitHub पर देखें↗

    Zed is an AI-native, high-performance code editor designed for extreme responsiveness and keyboard-centric workflows. It functions as an extensible text processing workspace that integrates autonomous agents and predictive models directly into the development environment to automate complex engineering tasks, refactoring, and code generation. The editor distinguishes itself through a GPU-accelerated rendering pipeline and an asynchronous multi-threaded architecture that ensures low-latency interaction even with large-scale projects. It features built-in support for real-time, multi-user colla

    Bridge local development sessions with remote cloud-based intelligence providers to access sophisticated code completion and analysis capabilities.

    Rustgpuirust-langtext-editor
    GitHub पर देखें↗85,338
  • vllm-project/vllmvllm-project का अवतार

    vllm-project/vllm

    83,048GitHub पर देखें↗

    vLLM is a high-throughput inference engine designed for the efficient serving and execution of large language models. It functions as a production-ready distributed model server, providing standard API protocols for online serving while also supporting offline batch processing. The system is built to maximize token generation speed and memory efficiency, enabling both large-scale cloud deployments and local execution on personal hardware. The project distinguishes itself through advanced memory management and request scheduling techniques, most notably its use of non-contiguous key-value cach

    Dynamically inserts new sequences into active inference batches to maximize hardware utilization.

    Pythonamdblackwellcuda
    GitHub पर देखें↗83,048
  • paddlepaddle/paddleocrPaddlePaddle का अवतार

    PaddlePaddle/PaddleOCR

    82,412GitHub पर देखें↗

    PaddleOCR is a comprehensive optical character recognition framework designed for detecting and transcribing text from images and documents into structured, machine-readable formats. It provides a modular computer vision pipeline that decouples image preprocessing, text detection, and character recognition into independent, configurable stages. This architecture supports automated document digitization and multilingual text recognition, capable of identifying text in over one hundred languages across diverse environments ranging from scanned documents to industrial scenes. The framework disti

    Executes neural network models on high-performance runtimes across CPUs, GPUs, and specialized hardware accelerators.

    Pythonai4sciencechineseocrdocument-parsing
    GitHub पर देखें↗82,412
  • nomic-ai/gpt4allnomic-ai का अवतार

    nomic-ai/gpt4all

    77,375GitHub पर देखें↗

    GPT4All is a cross-platform runtime environment designed to execute large language models directly on local consumer hardware. By leveraging an optimized C++ inference backend, it enables private, offline AI interactions without requiring an internet connection or external cloud services. The project provides a comprehensive ecosystem for managing the entire model lifecycle, including discovery, downloading, and configuration of local weights. What distinguishes the platform is its integrated retrieval-augmented generation engine, which allows users to index local documents into semantic vect

    Executes quantized language models using optimized C++ tensor computation libraries for local CPU and GPU hardware.

    C++ai-chatllm-inference
    GitHub पर देखें↗77,375
  • openhands/openhandsOpenHands का अवतार

    OpenHands/OpenHands

    77,330GitHub पर देखें↗

    OpenHands is an autonomous agent framework designed for software engineering workflows. It provides a modular platform for orchestrating AI agents that reason, plan, and execute tasks within isolated, containerized development environments. By integrating with standard version control and development tools, the system enables agents to autonomously navigate codebases, implement features, and resolve issues through iterative reasoning and tool execution. The platform distinguishes itself through a model-agnostic orchestrator that connects diverse language models to a unified tool registry. It

    Routes requests to different language models based on performance, cost, or capability requirements to optimize agent execution.

    Pythonagentartificial-intelligencechatgpt
    GitHub पर देखें↗77,330
  • twitter/the-algorithmtwitter का अवतार

    twitter/the-algorithm

    73,422GitHub पर देखें↗

    The algorithm is a distributed recommendation engine pipeline designed to construct and serve personalized content timelines. It functions as a multi-stage orchestration layer that aggregates candidate content from diverse social graphs and high-dimensional embedding spaces, processing user interaction data to deliver a unified, ranked experience. The system utilizes a high-performance machine learning serving infrastructure to execute deep learning models that predict engagement probabilities in real-time. It distinguishes itself through a hybrid retrieval strategy that combines graph-traver

    Deploys predictive models to score content relevance and user engagement probabilities in real-time.

    Scala
    GitHub पर देखें↗73,422
  • compvis/stable-diffusionCompVis का अवतार

    CompVis/stable-diffusion

    73,125GitHub पर देखें↗

    Stable Diffusion is a generative machine learning pipeline that synthesizes high-resolution visual content by performing iterative denoising within a compressed latent space. By mapping natural language embeddings into pixel outputs through conditioned probabilistic processes, the framework enables the generation of images from text prompts and the transformation of existing visual inputs based on semantic instructions. The architecture utilizes a modular execution environment that decouples model loading, scheduler logic, and inference components to support diverse hardware configurations. I

    Coordinates model loading, hardware acceleration, and output processing to streamline production-ready inference.

    Jupyter Notebook
    GitHub पर देखें↗73,125
  • josephmisiti/awesome-machine-learningjosephmisiti का अवतार

    josephmisiti/awesome-machine-learning

    72,867GitHub पर देखें↗

    This project is a comprehensive, community-driven directory of machine learning resources, software libraries, and educational materials. It serves as a centralized knowledge base for developers and researchers, organizing tools and frameworks by their primary programming language and technical domain to simplify discovery across the artificial intelligence ecosystem. The collection distinguishes itself by providing a cross-language development index that spans diverse programming environments, including C, C++, Rust, Clojure, and Python. It covers a wide range of specialized capabilities, fr

    Facilitates collaborative model training and analytics across decentralized data sources using unified architectural frameworks.

    Python
    GitHub पर देखें↗72,867
  • hiyouga/llama-efficient-tuninghiyouga का अवतार

    hiyouga/LLaMA-Efficient-Tuning

    72,239GitHub पर देखें↗

    This project is a fine-tuning framework and training pipeline designed to optimize and adapt large language and vision models. It provides a specialized toolkit for parameter-efficient tuning and supervised learning, serving as both a trainer for multimodal models and a deployment tool for serving fine-tuned models via high-performance inference engines. The framework focuses on reducing memory and compute requirements by updating a small subset of model parameters. It supports a wide range of adaptation strategies, including vision-language model training to align text, image, video, and aud

    Serves fine-tuned models via APIs or user interfaces using high-performance inference engines.

    Python
    GitHub पर देखें↗72,239
पिछला123456…20अगला
  1. Home
  2. Artificial Intelligence & ML
  3. Machine Learning
  4. Infrastructure
  5. Model Inference and Serving

सब-टैग एक्सप्लोर करें

  • Engines, Runtimes & Servers7 सब-टैग्स
  • Inference Engines8 सब-टैग्सRuntime environments designed to execute pre-trained neural network models with optimized performance and efficiency.
  • Inference Optimization6 सब-टैग्सTechniques and configurations that enhance model execution speed, reduce memory usage, and improve computational efficiency during inference.
  • Local AI Deployment Platforms2 सब-टैग्सPlatforms for deploying and managing language model interfaces and data processing tasks on local hardware.
  • Model Integration & Pipelines4 सब-टैग्स
  • Request Routing & Gateways7 सब-टैग्स
  • Runtime Interfaces & Orchestration4 सब-टैग्स
  • Scikit-Learn Inference ServersServing runtimes that load and execute Scikit-Learn models for inference. **Distinct from Model Inference and Serving:** Distinct from Model Inference and Serving: specifically targets Scikit-Learn model execution, not general inference serving.