awesome-repositories.com
Blog
MCP
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetServeur MCPÀ proposNotre méthodologiePresse
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

34 dépôts

Awesome GitHub RepositoriesInference Optimization Utilities

Tools focused on post-training conversion, compilation, and hardware-specific acceleration for deployment-ready models.

Explore 34 awesome GitHub repositories matching artificial intelligence & ml · Inference Optimization Utilities. Refine with filters or upvote what's useful.

Awesome Inference Optimization Utilities GitHub Repositories

Trouvez les meilleurs dépôts grâce à l'IA.Nous recherchons les dépôts les plus pertinents grâce à l'IA.
  • unslothai/unslothAvatar de unslothai

    unslothai/unsloth

    66,628Voir sur GitHub↗

    Unsloth is a high-performance training and inference platform designed to optimize the lifecycle of large language and multimodal models. It provides a comprehensive engine for fine-tuning, executing, and managing models locally, with a focus on reducing memory consumption and increasing compute speed on consumer-grade hardware. The platform distinguishes itself through hand-optimized kernels and automated computational graph techniques that maximize hardware throughput. It supports advanced training methodologies, including reinforcement learning for reasoning and efficient adapter-based fin

    Exports custom model weights into standard file formats to ensure compatibility with local inference and production systems.

    Pythonagentdeepseekdeepseek-r1
    Voir sur GitHub↗66,628
  • keras-team/kerasAvatar de keras-team

    keras-team/keras

    64,094Voir sur GitHub↗

    Keras is a high-level deep learning framework designed for constructing and training neural networks through the composition of modular, functional layers. It serves as a comprehensive modeling toolkit that provides standardized procedures for defining, evaluating, and deploying complex architectures. By utilizing a directed acyclic graph approach, the framework allows users to build intricate models with multiple inputs, outputs, and shared layers, ensuring consistent numerical execution through functional state management. The project distinguishes itself as a multi-backend machine learning

    Applies hardware-specific tuning to model execution paths, significantly enhancing inference speed and throughput on diverse computing devices.

    Pythondata-sciencedeep-learningjax
    Voir sur GitHub↗64,094
  • ultralytics/yolov5Avatar de ultralytics

    ultralytics/yolov5

    57,528Voir sur GitHub↗

    YOLOv5 is a comprehensive computer vision framework designed for end-to-end deep learning, specializing in real-time object detection, image classification, and instance segmentation. It provides a unified toolkit that manages the entire lifecycle of a model, from initial dataset configuration and hyperparameter tuning to high-speed inference and deployment. The framework utilizes a modular neural architecture, allowing users to swap backbone and head components to tailor models for specific visual tasks. What distinguishes this project is its focus on production-ready deployment and model ef

    Translates trained models into standard industry formats to ensure compatibility across diverse hardware and deployment environments.

    Pythoncoremldeep-learningios
    Voir sur GitHub↗57,528
  • deepfakes/faceswapAvatar de deepfakes

    deepfakes/faceswap

    55,289Voir sur GitHub↗

    Faceswap is a comprehensive framework for automated media manipulation and neural face synthesis. It provides a modular pipeline that manages the entire lifecycle of facial feature extraction, deep learning model training, and image conversion. By coordinating complex computer vision workflows, the system enables users to map facial identities between source and destination datasets while maintaining structural alignment and lighting consistency across video frames. The project distinguishes itself through a highly extensible plugin-based architecture that handles hardware-accelerated process

    Converts trained models into inference-ready versions by calculating required layers and configuring swap parameters.

    Pythondeep-face-swapdeep-learningdeep-neural-networks
    Voir sur GitHub↗55,289
  • pytorchlightning/pytorch-lightningAvatar de PyTorchLightning

    PyTorchLightning/pytorch-lightning

    31,189Voir sur GitHub↗

    PyTorch Lightning is a high-level deep learning framework for PyTorch that automates training loops and removes repetitive engineering boilerplate. It functions as a structured pipeline for managing machine learning experiments, providing a distributed training orchestrator and tools for mixed-precision training. The framework decouples scientific model architecture from the engineering required for infrastructure and scaling. This separation allows the same model code to execute across CPUs, GPUs, or TPUs through a hardware-agnostic execution engine and a centralized trainer that manages the

    Converts trained models into standardized industry formats for compatibility and deployment in production environments.

    Python
    Voir sur GitHub↗31,189
  • facefusion/facefusionAvatar de facefusion

    facefusion/facefusion

    28,806Voir sur GitHub↗

    Facefusion is a modular framework designed for automated image and video manipulation, specializing in tasks such as face swapping, enhancement, and restoration. It functions as a computer vision processing pipeline that chains independent machine learning modules to perform complex transformations, including facial animation, age modification, and lip synchronization. The system is built to handle both real-time interactive feeds and large-scale batch processing tasks. The platform distinguishes itself through a highly extensible architecture that supports custom processing modules and inter

    Applies high-performance libraries to graphics hardware to reduce inference latency and increase throughput.

    Pythonaideep-fakedeepfake
    Voir sur GitHub↗28,806
  • svc-develop-team/so-vits-svcAvatar de svc-develop-team

    svc-develop-team/so-vits-svc

    28,097Voir sur GitHub↗

    This project is a singing voice conversion tool based on VITS generative modeling. It transforms the identity of a singing voice to a target speaker while preserving the original melody, lyrics, and intonation. The system distinguishes itself through hybrid voice synthesis, allowing for the blending of multiple speaker identities via linear model interpolation. It utilizes cluster-based feature retrieval to increase target voice similarity and employs a diffusion probabilistic model as a post-processor to remove electronic artifacts and improve vocal clarity. The software covers a broad rang

    Converts internal feature indices into a format compatible with external runtimes for cross-environment conversion.

    Python
    Voir sur GitHub↗28,097
  • heartexlabs/label-studioAvatar de heartexlabs

    heartexlabs/label-studio

    27,626Voir sur GitHub↗

    Label Studio est un outil d'étiquetage de données multi-types et un espace de travail d'annotation de données conçu pour préparer des jeux de données pour l'entraînement en apprentissage automatique. Il fonctionne comme un pipeline de données intégré au cloud qui importe des données brutes depuis le stockage, gère le processus d'annotation et exporte les étiquettes dans des formats standardisés. La plateforme dispose d'un framework d'intégration de modèles d'apprentissage automatique qui se connecte à des serveurs de modèles externes. Cela permet l'annotation assistée par modèle et l'apprentissage actif, permettant au système d'effectuer un pré-étiquetage et d'affiner les prédictions basées sur les commentaires humains. Le logiciel fournit des outils de gestion de projet pour organiser les jeux de données et assigner des tâches aux utilisateurs via un accès basé sur les rôles. Il prend en charge divers types de données et utilise des adaptateurs de stockage agnostiques au backend pour se connecter à des systèmes de fichiers locaux ou à des fournisseurs de stockage cloud. L'application peut être installée via une configuration manuelle ou des déploiements en un clic sur une infrastructure cloud.

    Converts labeled data into standardized formats compatible with various machine learning models for training.

    TypeScript
    Voir sur GitHub↗27,626
  • tzutalin/labelimgAvatar de tzutalin

    tzutalin/labelImg

    25,012Voir sur GitHub↗

    labelImg est un outil d'annotation d'images de bureau et un utilitaire de préparation de jeux de données utilisé pour créer des jeux de données étiquetés pour l'entraînement en vision par ordinateur. Il fournit une interface graphique pour dessiner des boîtes englobantes autour des objets dans les images et leur assigner des étiquettes de classe pour construire des données de vérité terrain pour les modèles d'apprentissage automatique. Le logiciel prend spécifiquement en charge le format d'annotation XML Pascal VOC, exportant les coordonnées des images et les noms de classe dans des structures XML ou texte standard. Il permet aux utilisateurs de charger des listes de classes prédéfinies à partir de fichiers texte pour standardiser la dénomination sur l'ensemble d'un projet. Au-delà de l'étiquetage initial, l'outil couvre les flux de travail d'annotation d'images, y compris la visualisation des annotations enregistrées et la vérification manuelle des jeux de données. Cela inclut la possibilité de marquer les images comme vérifiées ou difficiles pour maintenir la qualité du jeu de données.

    Converts human-created labels into standardized formats like Pascal VOC XML for model training.

    Python
    Voir sur GitHub↗25,012
  • baidu/paddleAvatar de baidu

    baidu/paddle

    23,959Voir sur GitHub↗

    Paddle is a deep learning framework designed for building, training, and deploying large-scale machine learning models. It incorporates a distributed training engine for optimizing performance across multiple chips and a model inference engine for transforming trained models into production-ready formats for cross-platform execution. The platform features a heterogeneous hardware abstraction and a standardized software stack that allows models to run across diverse hardware architectures through a common interface. It also includes a scientific computing library capable of solving complex dif

    Transforms trained models into optimized binaries to increase inference speed and reduce runtime overhead.

    C++
    Voir sur GitHub↗23,959
  • pytorch/examplesAvatar de pytorch

    pytorch/examples

    23,752Voir sur GitHub↗

    This repository serves as a comprehensive collection of reference implementations for the PyTorch machine learning library. It provides practical examples for building, training, and deploying deep learning models, functioning as a toolkit for developers to explore neural network architectures and training workflows. The project distinguishes itself by offering concrete demonstrations of complex machine learning operations, ranging from computer vision tasks like object detection and depth estimation to the training of large-scale transformer models. These examples illustrate how to implement

    Converts trained neural network models into optimized formats for efficient inference on specialized hardware.

    Python
    Voir sur GitHub↗23,752
  • mlc-ai/mlc-llmAvatar de mlc-ai

    mlc-ai/mlc-llm

    22,057Voir sur GitHub↗

    MLC LLM is a machine learning compiler and inference engine designed to execute large language models locally across diverse hardware platforms, including desktop, mobile, and web environments. By utilizing machine learning compilation, the project transforms high-level model definitions into specialized, hardware-specific binary libraries. This process optimizes model weights and generates compute kernels tailored to the unique memory and processing characteristics of target graphics and mobile hardware. The engine distinguishes itself by providing a unified runtime abstraction that enables

    Transforms and optimizes model weights into specialized binary libraries for efficient execution across diverse hardware backends.

    Pythonlanguage-modelllmmachine-learning-compilation
    Voir sur GitHub↗22,057
  • onnx/onnxAvatar de onnx

    onnx/onnx

    20,358Voir sur GitHub↗

    ONNX is an open-source standard for machine learning interoperability that provides a unified format for representing neural network models. By defining a common set of operators and a standardized file structure, it enables models to be shared, exported, and executed consistently across different training frameworks and software ecosystems. The project functions as an intermediate representation layer that decouples model development from deployment. It utilizes a language-neutral binary serialization format to store model structures and weights, ensuring that computational graphs remain por

    Applies hardware-specific acceleration techniques and specialized runtime libraries to improve inference speed and efficiency.

    Pythonaiartificial-intelligencedeep-learning
    Voir sur GitHub↗20,358
  • huggingface/candleAvatar de huggingface

    huggingface/candle

    19,422Voir sur GitHub↗

    Candle is a minimalist machine learning framework and deep learning inference engine designed for the Rust programming language. It functions as a low-level tensor computation library, providing the necessary primitives for multi-dimensional array operations and mathematical transformations required to execute pre-trained neural network models. The framework distinguishes itself through a focus on memory efficiency and hardware utilization. It employs static-typed tensor operations to enforce shape validation and memory safety at compile time, while utilizing a lazy-loaded computational graph

    Provides ahead-of-time compilation of neural network models to optimize inference performance and reduce runtime latency.

    Rust
    Voir sur GitHub↗19,422
  • modelscope/funasrAvatar de modelscope

    modelscope/FunASR

    18,481Voir sur GitHub↗

    FunASR is an automatic speech recognition toolkit and multilingual speech-to-text engine designed to convert spoken audio into written text across more than fifty languages. It provides a framework for speaker diarization, an OpenAI-compatible transcription API for local server hosting, and speech models compatible with the ONNX format. The project distinguishes itself by supporting high-performance inference on edge hardware via self-contained binaries and portable model exports. It incorporates specialized capabilities for natural speech generation with adjustable timbre and emotional expre

    Converts trained models into universal industry formats for deployment via containerized runtimes.

    Pythonasraudiochinese
    Voir sur GitHub↗18,481
  • pytorch/visionAvatar de pytorch

    pytorch/vision

    17,743Voir sur GitHub↗

    This project is a comprehensive computer vision library for the PyTorch ecosystem, providing a standardized collection of neural network architectures, datasets, and high-performance transformation utilities. It serves as a foundational framework for building, training, and deploying deep learning models, offering a centralized model registry that allows developers to instantiate architectures with pre-trained weights for tasks such as image classification, object detection, and semantic segmentation. The library distinguishes itself through its modular approach to data and compute management

    Compiles machine learning models for specialized hardware accelerators to improve execution speed.

    Pythoncomputer-visionmachine-learning
    Voir sur GitHub↗17,743
  • lightning-ai/litgptAvatar de Lightning-AI

    Lightning-AI/litgpt

    13,431Voir sur GitHub↗

    LitGPT is a training and deployment framework for large language models, providing a suite of tools for pretraining, finetuning, quantizing, evaluating, and serving models within a production environment. It includes a dedicated training pipeline for adapting pretrained models to specific tasks, a quantization tool for reducing weight precision, and an inference server for hosting models via web interfaces. The framework supports high-performance model development through custom architecture implementation and the use of predefined recipes to standardize pretraining and finetuning. It enables

    Applies hardware-specific quantization and memory optimizations to improve the performance of production model inference.

    Python
    Voir sur GitHub↗13,431
  • microsoft/loraAvatar de microsoft

    microsoft/LoRA

    13,264Voir sur GitHub↗

    LoRA is a framework for parameter-efficient fine-tuning of large-scale neural networks. It functions by injecting trainable low-rank decomposition matrices into frozen model layers, allowing for task-specific adaptation while preserving the integrity of the original base model weights. The project distinguishes itself by enabling the direct merging of these trained low-rank matrices into primary model weights. This process eliminates additional computational overhead during inference, ensuring that adapted models maintain the same performance characteristics as the original architecture. Furt

    Merges task-specific adaptation matrices into primary model weights to eliminate inference overhead and minimize checkpoint storage size.

    Pythonadaptationdebertadeep-learning
    Voir sur GitHub↗13,264
  • paddlepaddle/paddleformersAvatar de PaddlePaddle

    PaddlePaddle/PaddleFormers

    12,981Voir sur GitHub↗

    PaddleFormers is a framework for the training, fine-tuning, and deployment of large language models. It provides a full lifecycle pipeline for executing large-scale model training and applying adaptation methods to align models with specialized tasks. The project focuses on scaling model operations through distributed training and hardware accelerator integration. It employs pipeline parallelism and mixed-precision training to manage memory and increase throughput across multiple hardware devices. The library includes a curated model zoo for serving pre-trained architectures and tools for pr

    Supports converting trained model weights into standardized industry formats for compatibility with external deployment engines.

    Pythonmodel
    Voir sur GitHub↗12,981
  • nvidia/tensorrt-llmAvatar de NVIDIA

    NVIDIA/TensorRT-LLM

    12,913Voir sur GitHub↗

    TensorRT-LLM is a platform and toolkit designed for compiling, optimizing, and serving transformer-based models on accelerated hardware. It functions as a framework that transforms machine learning models into efficient execution graphs, providing an engine to refine these models for specific hardware to maximize throughput and minimize latency during text generation. The project distinguishes itself through advanced execution strategies that manage the entire inference pipeline. It utilizes kernel-level fusion and static graph execution to optimize mathematical operations and computational f

    Transforms machine learning models into highly efficient execution graphs for accelerated text generation.

    Pythonblackwellcudallm-serving
    Voir sur GitHub↗12,913
Préc.12Suivant
  1. Home
  2. Artificial Intelligence & ML
  3. Machine Learning
  4. Infrastructure
  5. Optimization & Inference
  6. Serving & Runtime
  7. Inference Optimization Utilities

Explorer les sous-tags

  • Inference Optimization ToolsUtilities that apply hardware-specific optimizations to improve the performance of machine learning model inference.
  • Model Compilation1 sous-tagTools that transform trained machine learning models into optimized versions specifically prepared for efficient inference execution.
  • Model Export Formats2 sous-tagsUtilities for converting trained machine learning models into standard industry formats for compatibility and deployment.