awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektMCP-ServerÜber unsRanking-MethodikPresse
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
dusty-nv avatar

dusty-nv/jetson-inference

0
View on GitHub↗
developer.nvidia.com/embedded/twodaystoademo↗

Jetson Inference

jetson-inference is a set of libraries and tools for executing optimized deep learning models on embedded GPU hardware. Its primary purpose is to enable real-time computer vision and AI inference at the edge with low latency and high throughput.

The project distinguishes itself through high-performance streaming analytics and the ability to execute concurrent AI pipelines on auto-grade silicon. It provides specialized support for multi-sensor stream processing, utilizing zero-copy data transport to load camera frames directly into GPU memory.

The codebase covers a broad surface of capabilities, including real-time video analytics, object detection and tracking, and image segmentation. It also integrates hardware-accelerated decoding and TensorRT-based inference to optimize model execution on embedded platforms.

The project provides a TensorRT inference wrapper and an embedded vision SDK to facilitate the deployment of neural network primitives.

Features

  • Deep Learning Inference Engines - Provides a high-performance runtime engine for executing deep learning models on embedded GPU hardware via TensorRT.
  • GPU Accelerated Computer Vision - Uses GPU-optimized libraries to accelerate real-time image processing, depth estimation, and pose tracking.
  • Inference Execution - Executes optimized deep learning models on specialized GPU hardware to produce fast, accurate predictions.
  • Computer Vision Platforms - Provides a comprehensive environment for developing and deploying real-time computer vision applications on embedded hardware.

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Start searching with AI
8,734 Stars·3,089 Forks·C++·mit·14 Aufrufe
  • Edge AI Model Deployment - Running optimized deep learning models on embedded GPU hardware for real-time computer vision and robotics.
  • GPU Performance Profilers - Analyzes and debugs GPU-accelerated workloads to optimize AI, graphics, and compute performance.
  • AI Hosting Platforms - Deploys pretrained or customized AI models as GPU-accelerated containers using industry-standard APIs.
  • AI Model Integrations - Incorporates curated SDKs and pre-trained models to accelerate the addition of AI capabilities to applications.
  • AI Workload Orchestration - Provides specialized libraries for AI, mathematics, and data science to accelerate complex computational workloads.
  • Cross-Format Model Importers - Converts models from PyTorch, Hugging Face, and ONNX formats into high-performance inference engines.
  • Computer Vision Pipelines - Constructs real-time data processing workflows to streamline data movement from sensors to AI inference.
  • Computer Vision Libraries - Provides a library for executing optimized neural network primitives and computer vision tasks on edge devices.
  • Vision Pipeline Orchestrators - Develops streaming pipelines that ingest videos and preprocess frames for optimized vision AI models.
  • Object Detection - Localizes objects using 2D bounding boxes to provide a front end for pose estimation.
  • Gesture Recognition Libraries - Identifies common hand gestures such as waving or thumbs-up from real-time visual streams.
  • Python Bindings - Provides low-level Python bindings and core runtime functionalities for direct CUDA platform interaction.
  • GPU-Accelerated Inference - Accelerates the inference phase of machine learning models for image, video, and audio data on GPUs.
  • Robotics Pipeline Acceleration - Optimizes data transport and GPU resource utilization across robotics graphs using hardware-accelerated modules.
  • GPU Kernel Implementations - Executes asynchronous, fine-grained data movements initiated directly by GPU threads to eliminate CPU overhead.
  • Image Classification - Implements deep learning models like ResNet and VGG to identify objects and labels within images.
  • Multi-Stage Inference Pipelines - Links multiple models and preprocessing steps into a single execution graph for complex vision and audio workflows.
  • Inference API Servers - Exposes guardrailed inference through a standalone API server, Docker containers, or production microservices.
  • Inference Optimizations - Transforms neural network models to reduce latency and increase throughput for production deployment.
  • Model Compilation Memory Optimization - Manages memory allocation to enable the deployment of large foundation models on resource-constrained edge devices.
  • Model Performance Optimization - Uses accelerated engines to reduce response latency and increase throughput for specific GPU hardware.
  • Model Quantization - Converts high-precision checkpoints into quantized engines to reduce VRAM usage and increase speed.
  • Primitive Accelerators - Executes highly tuned GPU-accelerated routines for convolution, pooling, normalization, and activation layers.
  • Pose Estimation - Recognizes and tracks anatomical points on the human body within images or video streams.
  • Weight Quantization - Reduces model precision using FP8 and INT4 formats to lower memory usage and accelerate execution.
  • Neural Network Compression - Applies quantization, pruning, sparsity, and distillation to reduce model size and increase execution efficiency.
  • Robotics Perception Acceleration - Runs hardware-accelerated packages for high-performance perception and localization in robotic systems.
  • Concurrent Model Execution - Executes multiple deep learning inference streams simultaneously on auto-grade silicon for real-time tasks.
  • Ensemble Inference Pipelines - Links multiple models and pre- or post-processing steps into a single ensemble to handle complex inference workflows.
  • Video Analytics Pipelines - Provides pipelines for real-time object detection, tracking, and segmentation on live video streams.
  • Action Recognition - Analyzes video sequences to identify and classify specific human activities or behaviors over time.
  • Zero-Copy Framework Integrations - Shares data with deep learning frameworks via zero-copy interfaces to eliminate expensive memory transfers.
  • Neural Network Acceleration - NVIDIA provides optimized CUDA kernels for triangle attention and triangle multiplication to speed up processing of 3D data.
  • Device Management and Deployment - Deploys, scales, and updates AI applications and system software over the air across a fleet of edge devices.
  • GPU Acceleration - Provides a comprehensive suite of compilers and runtime libraries for building high-performance GPU-accelerated applications.
  • Deep Learning Acceleration - NVIDIA runs highly tuned GPU-accelerated routines for convolution, attention, matmul, pooling, and normalization.
  • ROS Libraries and Tools - Provides inference nodes to incorporate deep learning capabilities into Robot Operating System (ROS/ROS2) environments.
  • Sensor - NVIDIA handles high-bandwidth data from diverse sensors over Ethernet to enable real-time AI processing.
  • Sensor - Cleans, filters, and transforms raw sensor data into structured formats using GPU-accelerated libraries.
  • GPU-Accelerated Data Streams - Implements high-throughput, low-latency data streaming to share GPU data between systems.
  • Stream Analytics Processing - Analyzes concurrent video, audio, and image data using a streaming analytics toolkit for real-time understanding.
  • Image Buffer Sharing - Implements zero-copy data transport to move camera frames directly into GPU memory without duplication.
  • Shared Memory Transports - Implements zero-copy memory transport to share data buffers between libraries without expensive CPU-to-GPU transfers.
  • Microservice Infrastructure - Packages high-performance AI inference as secure, reliable microservices for deployment across clouds and data centers.
  • Hardware-Accelerated Decoders - Utilizes on-chip codecs to decompress video and image streams directly into GPU memory for real-time processing.
  • Real-Time Video Analysis - Builds vision applications that perform real-time video analytics with accelerated inference and object tracking.
  • Hardware-Accelerated Video Pipelines - Implements a framework for hardware-accelerated decoding, encoding, and processing of video and audio streams.
  • Multimedia Processing - Provides low-level hardware access to cameras and video processing for high-performance multimedia pipelines.
  • Stream Decoding - NVIDIA decompresses video from multiple popular codecs using on-chip hardware acceleration.
  • Stream Encoding - NVIDIA compresses video into various formats including H.264, HEVC, and AV1 using dedicated hardware.
  • Video Encoding and Decoding - NVIDIA accelerates video encoding and decoding using hardware-specific APIs on Windows and Linux.
  • Edge AI Perception Toolkits - Provides tools for integrating camera sensors and AI models into robotics and autonomous systems.
  • Sensor Processing - Builds high-performance sensor-processing pipelines using zero-copy data transport directly into GPU memory.
  • GPU Computations - Leverages parallel processing power on GPUs to execute computationally intensive tasks through Python applications.
  • Modular Camera Backends - Decouples sensor ingestion from inference logic to support diverse camera inputs across different hardware platforms.
  • GPU Memory Orchestration - Manages data transfers between GPUs using CPU-based operations and CUDA streams.
  • Unified Memory Managers - Manages low-level memory allocation and access between host and device to simplify GPU-accelerated development.
  • Zero-Copy Mechanisms - Uses columnar memory formats and zero-copy interfaces to minimize data transfer overhead between CPU and GPU.
  • Inference Performance Monitoring - Provides detailed observability metrics and Helm charts to monitor and scale AI inference microservices.
  • 3D Pose Reconstruction - Tracks skeletal movement from a single camera to reconstruct full-body 3D animations without markers.
  • AI Model Orchestration - Manages and executes both local and cloud AI models on edge devices for autonomous, multimodal applications.
  • Throughput Optimizations - Increases inference throughput using custom attention kernels, in-flight batching, and paged KV caching.
  • Cross-Camera Tracking - Maintains unique object identities across a network of multiple cameras to handle occlusions.
  • Object Pose Estimations - Tracks the 6D pose of novel objects using foundation models to determine exact position and orientation.
  • Monocular Depth Estimators - Predicts spatial depth from a single camera lens using monocular depth estimation algorithms.
  • Depth Estimation - Calculates distance and spatial geometry using both monocular and stereo depth estimation models.
  • Image Segmentation - Implements pixel-level classification to define precise shapes and boundaries of objects in images.
  • Inference Speed Profiling - Profiles model performance and analyzes execution timing to tune inference speed and efficiency.
  • Driving Scenario Generation - Produces photorealistic world variations from text prompts and spatial controls to expand driving datasets.
  • Generative AI Model Serving - Distributes generative AI inference workloads across GPU fleets using intelligent resource scheduling and request routing.
  • Generative AI Development - Provides an ecosystem of tools for developing conversational agents, copilots, and generative AI search engines.
  • GPU Acceleration - NVIDIA executes existing scikit-learn, UMAP, or HDBSCAN code on GPUs without requiring modifications to the source code.
  • Tile-Based Kernel Authoring - Implements a tile programming model in C++ and Python to manage high-performance data movement across GPU threads.
  • Multi-Node Inference Scaling - NVIDIA deploys large models across multiple GPUs and nodes using pipeline parallelism to handle models exceeding single-GPU memory.
  • Inference Acceleration - Reduces latency and increases throughput for large language model execution using a simplified API.
  • Inference Scaling Frameworks - NVIDIA integrates with Kubernetes and cloud orchestration environments to deploy and scale deep learning models across clusters.
  • Just-In-Time Kernel Compilers - Translates Python functions into optimized CUDA kernels at runtime for fine-grained thread control.
  • Large Language Model Serving - Provides high-speed inference and serving for large language models and vision language models.
  • Large-Scale Training Frameworks - Implements data and model parallelism for foundational scale models using a GPU-accelerated distributed framework.
  • Vision AI Agents - Develops intelligent vision applications that process visual data to automate tasks or monitor environments.
  • High-Performance AI Inference - Executes large language models with a modular runtime to maximize throughput on GPU hardware.
  • GPU Training Accelerators - Executes collective communication operations to distribute large models across multiple GPUs for faster training.
  • Mixed Precision Training - Reduces training time using multi-GPU distribution and mixed-precision floating-point computations.
  • Vision Model Fine-Tuning - Adapts pre-trained vision backbones and foundation models using domain-specific data and natural language prompts.
  • LLM Serving Architectures - Implements high-performance serving architectures for large and vision language models.
  • Generative AI Models - NVIDIA runs large language models and vision transformers on embedded hardware to enable real-time AI in robotics and computer vision.
  • Data Preprocessing - Decodes and augments images, videos, and speech in parallel with training to eliminate loading bottlenecks.
  • Automated Image Labeling - Automatically generates object detection and segmentation masks using AI-driven prompts and descriptors.
  • Multi-Framework Model Serving - Serves models from multiple frameworks across diverse hardware accelerators and CPUs using optimized configurations.
  • Multi-Physics Simulations - Calculates multi-physics behaviors for robotics and digital twins using GPU-accelerated engines.
  • Multimodal Analysis Tools - Combines image and video data with text prompts to perform feature extraction and segmentation.
  • Neural Network Design Frameworks - Provides an integrated environment for the structural design and development of deep neural networks for inference.
  • Real-Time Speech Processing - Develops customized, real-time speech applications using GPU-accelerated processing pipelines.
  • Multi-Camera Tiled Rendering - Consolidates multi-camera input into a single image via tiled rendering for real-time agent data.
  • Speech-to-Text Conversions - Transcribes spoken language into text with multi-language support and optimized memory footprints for on-device use.
  • Standardized AI Component Abstractions - Executes deep learning workloads using standardized programming models to ensure portability between cloud and embedded hardware.
  • Tensor Data Representations - Converts diverse 3D and multimedia formats into a consistent tensor representation for AI training and inference.
  • Autonomous Vehicle Dataset Curation - Scales the labeling and curation of autonomous vehicle datasets using integrated cloud hardware and enterprise software.
  • Training Data Generation - Produces augmented datasets by randomizing scene attributes like lighting and color to improve AI model robustness.
  • Transfer Learning - Adapts pre-trained models to specific platforms to optimize inference throughput.
  • Agentic Visual Reasoning - Builds intelligent systems that utilize computer vision and real-time visual reasoning to interact with the physical world.
  • Synthetic Video Generators - Generates synthetic single and multiview videos based on vehicle data to accelerate training scenarios.
  • Video Object Tracking - Follows objects across sequential video frames using optical flow to optimize GPU usage.
  • 3D Reconstruction - Converts RGB-D or lidar data into dense 3D maps and temporal costmaps for navigation.
  • Autonomous Driving - Combines reconstructed scenes with traffic and policy models for scalable closed-loop testing of self-driving systems.
  • Large Language Model Deployments - Hosts a wide variety of large language models via standardized microservices.
  • Model Evaluation and Benchmarking - Runs model benchmarks across local machines, HPC clusters, or cloud platforms using a unified interface.
  • Production Traffic Scaling - NVIDIA serves optimized models using dynamic batching, concurrent execution, and model ensembling to handle production traffic.
  • Robotics Simulators - Creates physically based virtual environments for robotics testing using rigid body and vehicle dynamics.
  • Simulation Environments - Reconstructs real-world data into interactive simulations to test autonomous driving workflows.
  • Synthetic Data Generation - Generates synthetic images and videos to expand training datasets and enhance model robustness.
  • World Models & Simulation - Generates high-fidelity 3D environments and sensor data to test autonomous systems against rare environmental conditions.
  • Robotic Sensor Simulation - Simulates perception hardware output, such as lidar and depth cameras, using GPU-accelerated rendering.
  • Physics Simulation - Provides a GPU-accelerated physics engine to calculate interactions for robotic systems.
  • Robotics Simulators - Provides virtual environments to train and validate robotic behaviors before deployment to physical hardware.
  • Multi-Stage Matrix Optimization - Executes multi-stage matrix-matrix multiplications with fusion and tuning to maximize hardware performance.
  • Collective GPU Communication - NVIDIA executes collective communication routines like all-reduce and broadcast to share data across multiple GPUs and nodes.
  • Storage Throughput Optimizers - NVIDIA bypasses CPU bounce buffers to move data directly between storage and GPU memory.
  • Dataset Preparation Tools - Ingests and converts raw data into optimized formats using pipeline management and automated labeling.
  • GPUDirect Storage - NVIDIA moves sensor data directly into GPU memory using high-speed capture cards to minimize ingestion latency.
  • Model-Assisted Labelers - Runs deep learning models to automatically label datasets with GPU-accelerated pre- and post-processing.
  • Model Weight Conversions - Translates model weights between different formats to ensure interoperability between training frameworks and inference engines.
  • GPU State Inspection - NVIDIA controls execution via breakpoints and single-stepping to inspect variables, registers, and GPU state.
  • Robotics System Integrations - Supports custom ROS2 messages and URDF formats to enable standalone scripting and manual control of simulations.
  • Automotive AI Deployment - Provides a full-stack platform to build and run scalable, real-time AI applications for automotive production.
  • AI Deployment Containers - Runs specialized AI functions using user-provided containers, models, and Helm charts.
  • LLM Inference Optimization - Accelerates large language models through prefix caching, key-value caching, and disaggregated serving.
  • Background Removal Tools - Separates primary subjects from their background for isolation or replacement.
  • Cloud Native Development Tools - Employs containers, Kubernetes, and microservices to create scalable AI applications bridging cloud and edge.
  • Cloud Native Infrastructure - Uses containerized development and Kubernetes to scale edge AI within cloud-native infrastructure.
  • Cloud Native GPU Orchestration - Scales compute workloads across on-premises, private, and public cloud resource clusters using GPU orchestration.
  • Deployment Orchestration - Standardizes the training and deployment workflow across edge and cloud environments with automated tuning.
  • Media Processing Scaling - Scales image and signal processing workloads across multiple GPUs to increase throughput.
  • GPU Container Toolkits - Configures container runtimes to enable hardware-accelerated applications to run inside portable containers.
  • Inference Engine Compilers - Creates lightweight, cross-OS and cross-GPU portable inference engines directly on target hardware.
  • GPU Resource Automation - Manages the software required to expose GPUs on Kubernetes to improve performance and utilization.
  • Kubernetes Deployment Management - Coordinates the startup ordering and scaling of interdependent inference components on Kubernetes.
  • Model Conversion - Parses models from PyTorch, Hugging Face, and ONNX to generate optimized inference engines.
  • 3D Rendering Engines - Uses GPU-accelerated APIs to perform high-performance 3D rendering and UI display.
  • Hardware-Accelerated Ray Tracing - Implements a flexible pipeline for ray generation, intersection, and shading to optimize GPU ray tracing.
  • Volumetric Ray Tracing - Uses hierarchical algorithms to perform fast ray tracing for city-scale neural radiance fields.
  • Volumetric Rendering Engines - Accelerates the rendering of sparse volumetric data structures for real-time visualization of complex effects.
  • Image Processing - Applies rectification, color correction, filtering, and feature extraction algorithms to optimize image data.
  • Custom Sensor Data Pipelines - Constructs flexible processing graphs using custom operators to transform audio, image, and video data.
  • Image Processing - Performs high-performance image processing and transformations directly on the GPU.
  • Computer Vision Operator Acceleration - Executes a specialized set of high-performance computer vision operators on the GPU to reduce processing costs.
  • Video Object Segmentations - Runs models on live video feeds to isolate specific objects using real-time query points.
  • Video Dataset Processing - Processes video content using GPU-accelerated pipelines for splitting and sharding large files into training datasets.
  • Motion Vector Calculation - Calculates relative pixel motion between frames using dedicated GPU hardware to track object movement.
  • Real-Time Video Filtering - NVIDIA accelerates video processing for effects including AI green screens, background blur, and webcam denoising.
  • Stereo Vision Reconstruction - Generates depth maps using stereo matching with zero-shot generalization for unfamiliar scenes.
  • Hardware-in-the-Loop Simulators - Tests and verifies trained robot behaviors in high-fidelity physical environments before hardware deployment.
  • Robotics And Autonomous Systems - Provides tools for building robotic systems including motion, perception, and autonomous navigation.
  • SLAM Algorithms - Implements high-performance visual SLAM to track robot position and map environments in real-time.
  • Real-Time Sensor Fusion - Processes multimodal data from images, video, and lidar to extract real-time environmental metadata.
  • Vehicle Egomotion Tracking - Predicts a vehicle's pose by applying motion models to odometry and IMU measurements.
  • GPU Memory Diagnostics - NVIDIA identifies memory access violations and detects precise exceptions using integrated memory checking tools.
  • GPU Shared Memory Race Detection - Detects hazardous data access patterns where multiple threads access the same shared memory location.
  • Uninitialized Memory Detectors - Flags instances where device global memory is read before it has been initialized.
  • Remote GPU Memory Access - NVIDIA moves data between local or remote storage and GPU memory using a direct-memory access engine to bypass the CPU.
  • Kernel Fusion Operations - NVIDIA combines multiple memory-bound and compute-bound operations into single kernels to reduce memory overhead.
  • Signal Processing - Executes GPU-accelerated primitives for color conversion, filtering, and geometry transforms.
  • In-Kernel Execution - Performs linear algebra operations directly on the device side within CUDA kernels to reduce latency.
  • Linear Algebra - Performs vector and matrix calculations using hardware acceleration for dense linear algebra workloads.
  • Distributed NumPy Workflows - Executes NumPy API operations across multiple GPUs and nodes to handle large-scale numerical computing.
  • Application Performance Optimization - Analyzes execution traces and hardware metrics to identify bottlenecks and increase GPU code efficiency.
  • GPU API Call Tracing - Registers callbacks for specific CUDA Runtime and Driver API calls to monitor entry and exit points.
  • Distributed Monitoring Tools - Profiles communication patterns and reliability to debug multi-node scaling across distributed systems.
  • GPU Profilers - Captures detailed logs of GPU kernel executions and memory operations with normalized timestamps.
  • Hardware Monitoring Tools - Reports real-time telemetry including GPU utilization, temperatures, and power draw.
  • Memory Leak Detection - Identifies out-of-bounds accesses, misaligned memory reads, and memory leaks during runtime.
  • Mobile and Embedded AI - Deep learning inference tutorials and tools for NVIDIA Jetson.
  • Star-Verlauf

    Star-Verlauf für dusty-nv/jetson-inferenceStar-Verlauf für dusty-nv/jetson-inference

    Häufig gestellte Fragen

    Was macht dusty-nv/jetson-inference?

    jetson-inference is a set of libraries and tools for executing optimized deep learning models on embedded GPU hardware. Its primary purpose is to enable real-time computer vision and AI inference at the edge with low latency and high throughput.

    Was sind die Hauptfunktionen von dusty-nv/jetson-inference?

    Die Hauptfunktionen von dusty-nv/jetson-inference sind: Deep Learning Inference Engines, GPU Accelerated Computer Vision, Inference Execution, Computer Vision Platforms, Edge AI Model Deployment, GPU Performance Profilers, AI Hosting Platforms, AI Model Integrations.

    Welche Open-Source-Alternativen gibt es zu dusty-nv/jetson-inference?

    Open-Source-Alternativen zu dusty-nv/jetson-inference sind unter anderem: nvidia/isaac-gr00t. paddlepaddle/paddledetection — PaddleDetection is an object detection framework designed for the end-to-end development, training, and deployment of… wang-xinyu/tensorrtx — tensorrtx is a computer vision inference engine and model implementation library designed for graphics processor… pytorch/executorch — ExecuTorch is a lightweight C++ runtime for deploying PyTorch models on mobile, embedded, and edge hardware. It… hybridgroup/gocv — GoCV is a computer vision library and Go language binding for OpenCV. It serves as an image processing toolkit and… tingsongyu/pytorch_tutorial — This project is a comprehensive collection of educational examples and reference implementations for building vision…

    Open-Source-Alternativen zu Jetson Inference

    Ähnliche Open-Source-Projekte, sortiert nach der Anzahl der gemeinsamen Funktionen mit Jetson Inference.
    • nvidia/isaac-gr00tAvatar von NVIDIA

      NVIDIA/Isaac-GR00T

      6,222Auf GitHub ansehen↗
      Jupyter Notebook
      Auf GitHub ansehen↗6,222
    • paddlepaddle/paddledetectionAvatar von PaddlePaddle

      PaddlePaddle/PaddleDetection

      14,243Auf GitHub ansehen↗

      PaddleDetection is an object detection framework designed for the end-to-end development, training, and deployment of computer vision models. It provides a comprehensive library of modular neural network architectures and pipelines that support object detection, instance segmentation, and multi-object tracking tasks. The project distinguishes itself through a configuration-driven approach that decouples model components like backbones and heads, allowing for the flexible assembly of custom vision workflows. It incorporates advanced techniques such as anchor-free detection logic, joint detecti

      Pythonblazefacedeepsortdetr
      Auf GitHub ansehen↗14,243
    • wang-xinyu/tensorrtxAvatar von wang-xinyu

      wang-xinyu/tensorrtx

      7,802Auf GitHub ansehen↗

      tensorrtx is a computer vision inference engine and model implementation library designed for graphics processor acceleration. It provides a framework for optimizing deep learning models through a GPU inference optimizer, a deep learning model converter for transforming weights from frameworks like TensorFlow and PyTorch, and a custom plugin library to implement operations not natively supported by the TensorRT API. The project distinguishes itself through a comprehensive collection of pre-defined network implementations, ranging from various YOLO versions and DETR transformers for object det

      C++arcfacecrnndetr
      Auf GitHub ansehen↗7,802
    • pytorch/executorchAvatar von pytorch

      pytorch/executorch

      4,296Auf GitHub ansehen↗

      ExecuTorch is a lightweight C++ runtime for deploying PyTorch models on mobile, embedded, and edge hardware. It provides an ahead-of-time compilation pipeline that exports, quantizes, and lowers model graphs into compact serialized programs, then executes them through a minimal runtime with hardware acceleration and on-device large language model inference capabilities. The project distinguishes itself through a hardware accelerator delegate system that partitions model subgraphs and offloads computation to specialized backends including NPUs, GPUs, and DSPs from Apple, Arm, Intel, MediaTek,

      Pythondeep-learningembeddedgpu
      Auf GitHub ansehen↗4,296
    Alle 30 Alternativen zu Jetson Inference anzeigen→