awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

89 repository-uri

Awesome GitHub RepositoriesInference Engines

Runtime environments designed to execute pre-trained neural network models with optimized performance and efficiency.

Explore 89 awesome GitHub repositories matching artificial intelligence & ml · Inference Engines. Refine with filters or upvote what's useful.

Awesome Inference Engines GitHub Repositories

Găsește cele mai bune repo-uri cu AI.Vom căuta cele mai potrivite repository-uri folosind AI.
  • jmorganca/ollamaAvatar jmorganca

    jmorganca/ollama

    174,350Vezi pe GitHub↗

    Ollama is a cross-platform runtime for managing, serving, and executing large language models on local hardware. It functions as a model manager and orchestrator that allows for the downloading, updating, and organization of model weights and configurations to ensure private and offline inference. The system provides a local inference API and a RESTful interface for programmatic model lifecycle management and text generation. It utilizes a compiled C++ backend to handle tensor operations and memory management. To support various hardware configurations, the runtime employs dynamic GPU offloa

    Provides a deployment environment that runs quantized models on local hardware with API support.

    Go
    Vezi pe GitHub↗174,350
  • vllm-project/vllmAvatar vllm-project

    vllm-project/vllm

    83,048Vezi pe GitHub↗

    vLLM is a high-throughput inference engine designed for the efficient serving and execution of large language models. It functions as a production-ready distributed model server, providing standard API protocols for online serving while also supporting offline batch processing. The system is built to maximize token generation speed and memory efficiency, enabling both large-scale cloud deployments and local execution on personal hardware. The project distinguishes itself through advanced memory management and request scheduling techniques, most notably its use of non-contiguous key-value cach

    Maximizes token generation speed and memory efficiency when serving large language models to multiple concurrent users.

    Pythonamdblackwellcuda
    Vezi pe GitHub↗83,048
  • paddlepaddle/paddleocrAvatar PaddlePaddle

    PaddlePaddle/PaddleOCR

    82,412Vezi pe GitHub↗

    PaddleOCR is a comprehensive optical character recognition framework designed for detecting and transcribing text from images and documents into structured, machine-readable formats. It provides a modular computer vision pipeline that decouples image preprocessing, text detection, and character recognition into independent, configurable stages. This architecture supports automated document digitization and multilingual text recognition, capable of identifying text in over one hundred languages across diverse environments ranging from scanned documents to industrial scenes. The framework disti

    Abstracts execution logic to allow seamless model operation across diverse CPU, GPU, and mobile hardware backends.

    Pythonai4sciencechineseocrdocument-parsing
    Vezi pe GitHub↗82,412
  • nomic-ai/gpt4allAvatar nomic-ai

    nomic-ai/gpt4all

    77,375Vezi pe GitHub↗

    GPT4All is a cross-platform runtime environment designed to execute large language models directly on local consumer hardware. By leveraging an optimized C++ inference backend, it enables private, offline AI interactions without requiring an internet connection or external cloud services. The project provides a comprehensive ecosystem for managing the entire model lifecycle, including discovery, downloading, and configuration of local weights. What distinguishes the platform is its integrated retrieval-augmented generation engine, which allows users to index local documents into semantic vect

    Executes quantized language models using optimized C++ tensor computation libraries for local CPU and GPU hardware.

    C++ai-chatllm-inference
    Vezi pe GitHub↗77,375
  • unslothai/unslothAvatar unslothai

    unslothai/unsloth

    66,628Vezi pe GitHub↗

    Unsloth is a high-performance training and inference platform designed to optimize the lifecycle of large language and multimodal models. It provides a comprehensive engine for fine-tuning, executing, and managing models locally, with a focus on reducing memory consumption and increasing compute speed on consumer-grade hardware. The platform distinguishes itself through hand-optimized kernels and automated computational graph techniques that maximize hardware throughput. It supports advanced training methodologies, including reinforcement learning for reasoning and efficient adapter-based fin

    Powers local execution of quantized models while enabling API support and tool-calling capabilities for external software.

    Pythonagentdeepseekdeepseek-r1
    Vezi pe GitHub↗66,628
  • ultralytics/ultralyticsAvatar ultralytics

    ultralytics/ultralytics

    58,468Vezi pe GitHub↗

    Ultralytics is a comprehensive computer vision framework designed for training, validating, and deploying deep learning models across a wide range of visual recognition tasks. It provides a unified interface for core operations including object detection, instance segmentation, pose estimation, and image classification. By utilizing a modular architecture, the platform allows users to swap model components to balance inference speed and accuracy requirements for diverse applications. The framework distinguishes itself through its support for real-time processing and flexible deployment. It in

    Executes pre-trained models on various data streams using highly optimized runtime environments.

    Pythonclicomputer-visiondeep-learning
    Vezi pe GitHub↗58,468
  • ultralytics/yolov5Avatar ultralytics

    ultralytics/yolov5

    57,528Vezi pe GitHub↗

    YOLOv5 is a comprehensive computer vision framework designed for end-to-end deep learning, specializing in real-time object detection, image classification, and instance segmentation. It provides a unified toolkit that manages the entire lifecycle of a model, from initial dataset configuration and hyperparameter tuning to high-speed inference and deployment. The framework utilizes a modular neural architecture, allowing users to swap backbone and head components to tailor models for specific visual tasks. What distinguishes this project is its focus on production-ready deployment and model ef

    Runs real-time object detection tasks by applying deep learning models to image and video streams.

    Pythoncoremldeep-learningios
    Vezi pe GitHub↗57,528
  • facebookresearch/segment-anythingAvatar facebookresearch

    facebookresearch/segment-anything

    54,353Vezi pe GitHub↗

    This project provides a deep learning architecture designed to identify and isolate distinct objects within images by generating precise pixel-level masks. It functions as a browser-based inference engine, enabling the execution of complex machine learning models directly within web environments without requiring server-side processing. The system distinguishes itself by utilizing hardware-accelerated execution and parallel processing to achieve real-time segmentation speeds. It supports prompt-based mask decoding, allowing users to generate spatial masks by providing specific points or boxes

    Leverages cross-platform runtime environments to execute pre-compiled models with consistent performance across varying hardware configurations.

    Jupyter Notebook
    Vezi pe GitHub↗54,353
  • ggerganov/whisper.cppAvatar ggerganov

    ggerganov/whisper.cpp

    50,791Vezi pe GitHub↗

    whisper.cpp is a C++ implementation of the Whisper speech-to-text model, serving as a lightweight machine learning inference engine and quantized runtime. It provides high-performance automatic speech recognition and real-time audio transcription without requiring a Python environment. The project utilizes model quantization to reduce memory usage and increase inference speed on local hardware. It incorporates hardware acceleration to optimize processing speed across different processors. The system covers audio processing capabilities including voice activity detection, speaker diarization,

    Provides a lightweight inference engine implemented in C to minimize runtime overhead and dependencies.

    C++
    Vezi pe GitHub↗50,791
  • microsoft/vibevoiceAvatar microsoft

    microsoft/VibeVoice

    49,394Vezi pe GitHub↗

    VibeVoice is a generative artificial intelligence platform designed for text-to-speech synthesis. It functions as a neural audio generation framework that converts written text into natural-sounding spoken audio, specifically engineered to maintain consistent vocal characteristics and narrative prosody across extended passages of content. The system distinguishes itself through its ability to generate long-form conversational speech while preserving speaker identity and linguistic content. By utilizing latent space disentanglement, the model separates speaker traits from the input text, allow

    Supports real-time streaming inference by processing audio generation in sequential chunks to minimize latency.

    Python
    Vezi pe GitHub↗49,394
  • lm-sys/fastchatAvatar lm-sys

    lm-sys/FastChat

    39,472Vezi pe GitHub↗

    FastChat is a training and serving platform for large language models that provides an integrated toolkit for fine-tuning, hosting, and benchmarking chatbots. It functions as an inference server capable of hosting multiple models and exposing them via a standardized API for chat applications. The platform distinguishes itself through a distributed model controller that manages worker nodes and routes requests across a hardware-agnostic inference layer supporting various accelerators. It includes a dedicated evaluation framework for assessing model quality using automated judges, multi-turn di

    Provides an abstraction layer that decouples model execution logic from specific GPU, CPU, or NPU hardware backends.

    Python
    Vezi pe GitHub↗39,472
  • bvlc/caffeAvatar BVLC

    BVLC/caffe

    34,576Vezi pe GitHub↗

    Caffe is a high-performance deep learning framework designed for training and deploying deep neural networks. It functions as a machine learning engine and a convolutional neural network library, providing a C++ backend to accelerate computations on both GPUs and CPUs. The system includes a specialized toolset for computer vision, enabling tasks such as object detection, semantic segmentation, and large-scale image retrieval. It supports the deployment of pre-trained models for image and scene recognition, as well as the ability to fine-tune neural network weights for specialized tasks. The

    Provides a high-performance C++ and CUDA backend to accelerate deep learning computations on GPUs and CPUs.

    C++deep-learningmachine-learningvision
    Vezi pe GitHub↗34,576
  • facebookresearch/detectron2Avatar facebookresearch

    facebookresearch/detectron2

    34,548Vezi pe GitHub↗

    Detectron2 is a PyTorch computer vision framework and visual recognition platform designed for training and deploying models for object detection, image segmentation, and visual recognition. It provides a research-oriented environment for training complex vision models with multi-GPU acceleration. The project includes a specialized object detection library for identifying and locating multiple objects via bounding boxes, as well as an image segmentation toolkit for creating pixel-level masks through instance, semantic, and panoptic segmentation. Additionally, it features a human pose estimati

    Executes pre-trained vision models on images, videos, or webcam feeds for real-time detection and segmentation.

    Python
    Vezi pe GitHub↗34,548
  • openbmb/voxcpmAvatar OpenBMB

    OpenBMB/VoxCPM

    29,985Vezi pe GitHub↗

    VoxCPM is a multilingual speech synthesis system and text-to-speech inference server. It functions as an AI voice cloning tool and a synthetic voice designer, capable of generating natural speech across global languages and regional dialects using a GPU-accelerated audio generator. The project features a speech model fine-tuning framework that supports both full parameter updates and low-rank adaptation for customizing voice characteristics. It enables high-fidelity voice cloning from reference audio, including cross-lingual voice transfer and acoustic environment mimicry, as well as the crea

    Implements a standardized runtime format that enables model execution across CUDA, MPS, and CPU backends.

    Pythonaudiodeeplearningminicpm
    Vezi pe GitHub↗29,985
  • sgl-project/sglangAvatar sgl-project

    sgl-project/sglang

    29,079Vezi pe GitHub↗

    Sglang is a high-performance inference engine and serving system designed for large language and multimodal models. It provides a programmable interface for orchestrating complex generation workflows, enabling developers to coordinate multi-turn dialogues, tool invocations, and reasoning chains through a domain-specific language. The platform is built to support production-scale deployments, offering an OpenAI-compatible API that allows for integration with existing application ecosystems. The system distinguishes itself through a disaggregated architecture that separates compute-intensive pr

    Provides a high-performance inference engine framework for serving large language models with complex workflow orchestration.

    Pythonattentionblackwellcuda
    Vezi pe GitHub↗29,079
  • mozilla-ocho/llamafileAvatar Mozilla-Ocho

    Mozilla-Ocho/llamafile

    25,090Vezi pe GitHub↗

    llamafile is a model bundler and local runtime that packages large language models and their execution logic into single, portable executable files. It provides a distribution format for zero-installation local execution, allowing users to run models on various operating systems without managing external library dependencies or environment configurations. The project differentiates itself by bundling model weights and the runtime into one self-extracting binary. This approach simplifies the distribution of AI models, as the combined file contains everything necessary to run the model immediat

    Integrates a lightweight C-based inference engine for minimal runtime overhead.

    C++
    Vezi pe GitHub↗25,090
  • anjok07/ultimatevocalremoverguiAvatar Anjok07

    Anjok07/ultimatevocalremovergui

    23,673Vezi pe GitHub↗

    Ultimate Vocal Remover is a desktop application designed for AI-driven audio source separation. It utilizes deep learning models to isolate vocals, drums, and other individual instruments from mixed audio files, providing a utility for professional production and creative editing workflows. The software distinguishes itself by leveraging GPU-accelerated tensor computation to perform complex signal processing tasks, significantly reducing the time required for high-fidelity audio extraction. It incorporates a modular plugin architecture that integrates external utilities to support a wide rang

    Executes pre-trained neural networks to perform complex pattern recognition and source separation on audio data.

    Pythonaudioinstrumentalkaraoke
    Vezi pe GitHub↗23,673
  • alexeyab/darknetAvatar AlexeyAB

    AlexeyAB/darknet

    22,159Vezi pe GitHub↗

    Darknet is a high-performance C-based inference engine and computer vision library designed for real-time object identification and localization. It serves as a neural network framework for training and deploying detection models using the YOLO architecture, providing a toolset for deep learning training and deployment. The project differentiates itself through a C and CUDA implementation that enables hardware acceleration for matrix multiplication and inference speed optimization. It provides a shared library interface for embedding detection capabilities into external applications and suppo

    Implements a high-performance inference engine in C and CUDA for minimal runtime overhead during image processing.

    C
    Vezi pe GitHub↗22,159
  • danielgatis/rembgAvatar danielgatis

    danielgatis/rembg

    21,911Vezi pe GitHub↗

    Rembg is a machine learning-based toolkit designed for automated image background removal and subject segmentation. It functions as a versatile engine that identifies and extracts subjects from images, supporting diverse input methods including individual files, directory-based batch processing, and live binary data streams. The project distinguishes itself through its flexible integration options, offering a command-line interface for local automation, a library for programmatic access, and an HTTP service for remote requests. It utilizes deep learning architectures to classify pixels and ge

    Executes pre-trained machine learning models using a cross-platform engine to ensure consistent performance across different hardware and operating systems.

    Pythonbackground-removalimage-processingpython
    Vezi pe GitHub↗21,911
  • dmlc/mxnetAvatar dmlc

    dmlc/mxnet

    20,812Vezi pe GitHub↗

    MXNet is a deep learning framework and distributed machine learning engine designed for training and deploying neural networks. It functions as a hardware-agnostic backend that allows for the development of deep learning models through a hybrid of symbolic and imperative programming. The system distinguishes itself through automatic distributed parallelism, which scales training workloads across multiple GPUs and machines. It features an extensible hardware backend interface that enables the integration of custom accelerators and proprietary libraries without modifying the core source code.

    Implements a hardware-agnostic backend that decouples model execution logic from specific hardware accelerators.

    C++
    Vezi pe GitHub↗20,812
Înapoi1234…5Înainte
  1. Home
  2. Artificial Intelligence & ML
  3. Machine Learning
  4. Infrastructure
  5. Model Inference and Serving
  6. Inference Engines

Explorează sub-etichetele

  • C++ Inference Backends2 sub-tag-uriHigh-performance tensor computation engines written in C++.
  • Computer Vision InferenceExecution of vision-based models using standard libraries for real-time object detection.
  • Deep LearningHigh-performance runtimes that execute neural network models across CPUs, GPUs, and specialized accelerators.
  • Hardware-Agnostic Inference LayersAbstraction layers that decouple model execution logic from specific hardware backends.
  • Local Inference RuntimesDeployment environments that run quantized models on local hardware with API support.
  • ONNX Runtime Inference2 sub-tag-uriExecuting models using the cross-platform ONNX runtime for consistent performance.
  • Request Schedulers1 sub-tagComponents that manage and prioritize incoming inference requests to optimize throughput and latency.
  • Streaming Inference ProcessorsExecution engines designed to process continuous streams of data using memory-efficient generators.