awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
NVIDIA avatar

NVIDIA/TensorRT

0
View on GitHub↗
13,076 stars·2,377 forks·C++·Apache-2.0·12 viewsdeveloper.nvidia.com/tensorrt↗

TensorRT

TensorRT is a deep learning inference engine and software development kit designed to optimize and deploy neural networks for high-performance execution on NVIDIA GPUs. It functions as a GPU acceleration framework that reduces latency and increases throughput for trained models during production deployment.

The toolkit imports models from the Open Neural Network Exchange format and transforms them into optimized engines. It utilizes graph-based model optimization, layer-fusion kernel generation, and precision-based quantization to convert floating point weights into lower precision formats.

The framework provides capabilities for hardware-specific engine serialization and supports the extension of inference capabilities through custom plugins for specialized neural network layers.

Features

  • Model Inference Accelerators - Transforms neural networks into high-performance engines to maximize execution speed on NVIDIA GPUs.
  • Cross-Format Model Importers - Imports model definitions from the ONNX format to prepare them for optimized GPU execution.
  • ONNX Model Importers - Parses Open Neural Network Exchange models to build internal representations for GPU optimization.
  • GPU Inference SDKs - Provides a comprehensive SDK for optimizing and deploying deep learning models on NVIDIA GPUs.
  • GPU Model Deployments - Enables the deployment of optimized deep learning models on NVIDIA GPU hardware accelerators.
  • GPU-Accelerated - Optimizes deep learning models for maximum throughput and low latency on GPU accelerators.
  • Deep Learning - Serves as a high-performance runtime environment that executes neural networks using NVIDIA GPU acceleration.
  • ONNX Engine Conversions - Converts models from the ONNX format into high-performance engines for NVIDIA GPU execution.
  • ONNX Model Optimizers - Imports ONNX models and transforms them into optimized engines for faster inference.
  • Hardware-Specific Model Optimizations - Compiles models into binary engines optimized for specific NVIDIA GPU architectures and memory limits.
  • Model Graph Optimizers - Provides graph-level optimizations by fusing layers and removing redundant operations to improve inference performance.
  • Neural Network Deployment - Provides the runtime and tools necessary to execute trained neural networks in production environments.
  • Precision Quantization - Converts floating point weights to lower precision formats like FP16 or INT8 to increase throughput.
  • GPU Acceleration - Provides a framework of tools to reduce latency and increase throughput for models deployed on GPUs.
  • Deep Learning Acceleration - Accelerates deep learning tensor operations and matrix multiplications on NVIDIA GPU hardware.
  • Custom Neural Network Layers - Allows for the implementation of specialized neural network layers via custom plugins.
  • Kernel Fusion Compilers - Generates fused kernels that combine multiple neural network layers to reduce memory bandwidth overhead.
  • Inference Capability Extensions - Allows adding specialized operations or layers to the runtime through custom plugin implementation.
  • Custom Operator Plugins - Supports the execution of custom neural network layers via external C++ plugin implementations.
  • AI & Machine Learning - High-performance inference on NVIDIA GPUs
  • Parallel Processing - High-performance inference library for NVIDIA GPUs.
  • Computation and Optimization - C++ library for high-performance inference on NVIDIA hardware.
  • Parallel and High-Performance Computing - High-performance inference library for NVIDIA GPUs.

Star history

Star history chart for nvidia/tensorrtStar history chart for nvidia/tensorrt

How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Open-source alternatives to TensorRT

Similar open-source projects, ranked by how many features they share with TensorRT.
  • dusty-nv/jetson-inferencedusty-nv avatar

    dusty-nv/jetson-inference

    8,734View on GitHub↗

    jetson-inference is a set of libraries and tools for executing optimized deep learning models on embedded GPU hardware. Its primary purpose is to enable real-time computer vision and AI inference at the edge with low latency and high throughput. The project distinguishes itself through high-performance streaming analytics and the ability to execute concurrent AI pipelines on auto-grade silicon. It provides specialized support for multi-sensor stream processing, utilizing zero-copy data transport to load camera frames directly into GPU memory. The codebase covers a broad surface of capabiliti

    C++caffecomputer-visiondeep-learning
    View on GitHub↗8,734
  • tencent/tnnTencent avatar

    Tencent/TNN

    4,641View on GitHub↗

    TNN is a deep learning inference framework designed to execute pre-trained neural networks across mobile, desktop, and server hardware. It functions as a hardware-accelerated runtime and model compression toolkit, providing a unified interface for deploying models in diverse environments. The framework includes an ONNX model converter to transform models from various training frameworks into a standardized internal format. It distinguishes itself through a combination of model compression tools—including weight quantization and static-code pruning—and a memory management system that reuses bu

    C++
    View on GitHub↗4,641
  • pytorch/executorchpytorch avatar

    pytorch/executorch

    4,296View on GitHub↗

    ExecuTorch is a lightweight C++ runtime for deploying PyTorch models on mobile, embedded, and edge hardware. It provides an ahead-of-time compilation pipeline that exports, quantizes, and lowers model graphs into compact serialized programs, then executes them through a minimal runtime with hardware acceleration and on-device large language model inference capabilities. The project distinguishes itself through a hardware accelerator delegate system that partitions model subgraphs and offloads computation to specialized backends including NPUs, GPUs, and DSPs from Apple, Arm, Intel, MediaTek,

    Pythondeep-learningembeddedgpu
    View on GitHub↗4,296
  • paddlepaddle/fastdeployPaddlePaddle avatar

    PaddlePaddle/FastDeploy

    3,700View on GitHub↗

    FastDeploy is a high-performance deployment framework for large language models, vision models, and multimodal models. It provides the infrastructure to launch model services that process combined image, video, and text inputs, exposing these capabilities through a standardized, OpenAI-compatible API for chat and text completions. The project distinguishes itself through advanced inference pipeline engineering and GPU optimization. It employs speculative decoding, tensor parallelism, and a disaggregated execution model that separates prefill and decode phases across different hardware resourc

    Pythonernieernie-45ernie-45-vl
    View on GitHub↗3,700
See all 30 alternatives to TensorRT→

Frequently asked questions

What does nvidia/tensorrt do?

TensorRT is a deep learning inference engine and software development kit designed to optimize and deploy neural networks for high-performance execution on NVIDIA GPUs. It functions as a GPU acceleration framework that reduces latency and increases throughput for trained models during production deployment.

What are the main features of nvidia/tensorrt?

The main features of nvidia/tensorrt are: Model Inference Accelerators, Cross-Format Model Importers, ONNX Model Importers, GPU Inference SDKs, GPU Model Deployments, GPU-Accelerated, Deep Learning, ONNX Engine Conversions.

What are some open-source alternatives to nvidia/tensorrt?

Open-source alternatives to nvidia/tensorrt include: dusty-nv/jetson-inference — jetson-inference is a set of libraries and tools for executing optimized deep learning models on embedded GPU… tencent/tnn — TNN is a deep learning inference framework designed to execute pre-trained neural networks across mobile, desktop, and… pytorch/executorch — ExecuTorch is a lightweight C++ runtime for deploying PyTorch models on mobile, embedded, and edge hardware. It… paddlepaddle/fastdeploy — FastDeploy is a high-performance deployment framework for large language models, vision models, and multimodal models.… nvidia/isaac-gr00t. tingsongyu/pytorch_tutorial — This project is a comprehensive collection of educational examples and reference implementations for building vision…