14 Repos
Capabilities for modifying input dimensions, batch sizes, and memory layouts during model execution.
Distinct from Dynamic Sync Shapes: None of the candidates address neural network tensor shapes; most relate to UI layouts or audio timbre.
Explore 14 awesome GitHub repositories matching artificial intelligence & ml · Dynamic Tensor Shapes. Refine with filters or upvote what's useful.
OpenVINO is an AI inference engine and model serving platform designed to execute optimized deep learning models across CPUs, GPUs, and NPUs through a unified API. It includes a model optimization toolkit for converting, quantizing, and compressing models from various frameworks, alongside a specialized generative AI runtime for large language models. The project distinguishes itself through a plugin-based hardware acceleration layer that maps neural network operations to vendor-specific drivers. It features advanced execution mechanisms such as continuous batching, speculative decoding, and
Allows modifying batch size and input shapes during execution to optimize the balance between throughput and latency.
This project is a structured learning curriculum and technical reference for mastering deep learning with TensorFlow. It provides a comprehensive guide for building, training, and deploying neural networks, combining theoretical fundamentals with practical implementation examples. The repository distinguishes itself by covering the end-to-end machine learning workflow, from low-level tensor mathematics and linear algebra to the creation of complex model architectures. It includes specific guidance on developing data pipelines for diverse data types, such as images, text, and time-series seque
Expands smaller tensor dimensions to match larger ones for compatible element-wise operations.
tensorrtx is a computer vision inference engine and model implementation library designed for graphics processor acceleration. It provides a framework for optimizing deep learning models through a GPU inference optimizer, a deep learning model converter for transforming weights from frameworks like TensorFlow and PyTorch, and a custom plugin library to implement operations not natively supported by the TensorRT API. The project distinguishes itself through a comprehensive collection of pre-defined network implementations, ranging from various YOLO versions and DETR transformers for object det
Provides capabilities for managing variable batch sizes and input dimensions through optimization profiles during model execution.
Flashlight ist eine eigenständige C++-Bibliothek für maschinelles Lernen und Tensor-Berechnungen, die zum Erstellen und Trainieren neuronaler Netze verwendet wird. Sie fungiert als umfassendes Framework für neuronale Netze und Engine für automatische Differenzierung und bietet Werkzeuge zur Konstruktion von Berechnungsgraphen und zur Berechnung von Gradienten via Backpropagation. Das Projekt dient als Framework für verteiltes Training und nutzt All-Reduce-Operationen zur Synchronisation von Gradienten und Parametern über mehrere Rechenknoten und Geräte hinweg. Es zeichnet sich durch eine tiefe Integration von leistungsstarker Tensor-Manipulation, nativer Interoperabilität mit Gerätespeichern und einem System zur Synchronisation von Gewichten über verteilte Worker aus, um das Training großskaliger Modelle zu beschleunigen. Das Framework deckt eine breite Palette an Deep-Learning-Funktionen ab, einschließlich modularer Schichtkomposition für den Entwurf komplexer Architekturen wie Residual-Blöcke und rekurrente Zellen. Es bietet umfangreiche Datenmanagement-Utilities für Ingestion und Prefetching sowie Serialisierungssysteme zur Persistierung von Modellzuständen. Zusätzlich enthält es eine Suite an Überwachungs- und Observability-Tools zur Verfolgung von Trainingsmetriken und zur Messung von Sequenzfehlern. Die Bibliothek ist in C++ implementiert.
Changes dimensions and permutes axes of tensors during model execution.
Pyrefly is a static type checker for Python that operates as a language server, delivering real-time diagnostics, completions, and navigation in any editor supporting the Language Server Protocol. It also performs static tensor shape analysis, using symbolic dimension variables and arithmetic to verify shape consistency in deep learning models without runtime execution. Beyond core type checking, Pyrefly supports gradual adoption workflows: it can generate a baseline of known errors so only new issues are reported, migrate configuration from other type checkers, and automatically suppress exi
Validates tensor shape consistency at compile time using symbolic arithmetic and shape rules for operations.
TileLang is a Python-embedded domain-specific language compiler that JIT-compiles and autotunes GPU kernels. It uses a tile-based DSL, automatic software pipelining, and parallel autotuning to generate optimized GPU kernels at runtime. It supports tensor core operations with Pythonic syntax, automatic memory management, and thread mapping. The compiler searches over tile sizes, thread counts, and scheduling policies, compiling and benchmarking candidates in parallel to find the fastest kernel. It also caches compiled binaries and tuning results to disk for reuse across sessions. TileLang inc
Queries supported tensor instruction shapes from the target architecture for instruction selection.
Dieses Projekt ist eine umfassende Lehrressource und ein Kurs zum Aufbau neuronaler Netze mit PyTorch. Es deckt die grundlegenden Bausteine des Deep Learning ab, einschließlich Tensor-Manipulation, automatischer Differenzierung und der Konstruktion modularer Komponenten für neuronale Netze. Das Repository dient als technischer Leitfaden für verschiedene spezialisierte Bereiche. Es bietet Implementierungsdetails für Computer-Vision-Aufgaben wie Bildklassifizierung, Objekterkennung und semantische Segmentierung sowie Workflows für die Verarbeitung natürlicher Sprache (NLP) mit Transformern, rekurrenten Netzen und generativen Modellen. Zudem enthält es eine Referenz für generative KI, mit Fokus auf die Synthese von Bildern mittels Diffusionsmodellen und adversarialen Netzwerken. Das Material erstreckt sich auf Modelloptimierung und Deployment-Pipelines. Es behandelt Techniken zur Reduzierung der Modellgröße und zur Erhöhung der Inferenzgeschwindigkeit durch Quantisierung und den Export von Modellen in Formate wie ONNX und TensorRT. Weitere Kompetenzbereiche umfassen Data Engineering für paralleles Laden, Modellevaluierung mittels benutzerdefinierter Metriken und das Deployment von Open-Source Large Language Models. Das Projekt wird primär als eine Reihe von Jupyter Notebooks bereitgestellt.
Changes tensor structures through operations such as reshaping, squeezing, permuting, and rotating axes.
ExecuTorch is a lightweight C++ runtime for deploying PyTorch models on mobile, embedded, and edge hardware. It provides an ahead-of-time compilation pipeline that exports, quantizes, and lowers model graphs into compact serialized programs, then executes them through a minimal runtime with hardware acceleration and on-device large language model inference capabilities. The project distinguishes itself through a hardware accelerator delegate system that partitions model subgraphs and offloads computation to specialized backends including NPUs, GPUs, and DSPs from Apple, Arm, Intel, MediaTek,
Handles tensor dimension changes between inference runs without requiring model recompilation.
PaddleRec ist eine Deep-Learning-Empfehlungsbibliothek und ein Framework für verteiltes Modelltraining auf Basis des PaddlePaddle-Frameworks. Es bietet eine Suite industrieller Algorithmen und Modelle für das User-Matching und das personalisierte Content-Ranking. Das Projekt enthält eine Empfehlungs-Inferenz-Engine für den Export und die Bereitstellung trainierter Modelle in Produktionsumgebungen für Online-Echtzeitanfragen. Es ermöglicht die Implementierung von Deep-Learning-Empfehlungsalgorithmen zur Verarbeitung massiver Verhaltensdatensätze. Das Framework deckt das Modelltraining im großen Maßstab über verteilte Rechencluster hinweg ab sowie die Entwicklung von Systemen zum Ranking von Elementen basierend auf persönlichen Präferenzen.
Supports dynamic tensor shapes to handle variable-length input sequences in recommendation ranking.
pytorch-summary ist eine Sammlung von Utilities für PyTorch-neuronale Netze, die dazu dienen, Modellzusammenfassungen zu generieren, Speicheranforderungen zu berechnen und Layer-für-Layer-Tensor-Shapes zu visualisieren. Es fungiert als Reporting-Tool, das detaillierte Aufschlüsselungen von Netzwerkschichten und Ausgabe-Shapes bereitstellt, um beim Debugging und der Inspektion von Modellen zu unterstützen. Das Projekt bietet spezialisierte Funktionen zur Schätzung des Gesamtspeicherverbrauchs von Forward- und Backward-Passes basierend auf Eingabedimensionen und Parameteranzahl. Es generiert menschenlesbare Visualisierungen von Modellstrukturen, um architektonische Entwürfe zu verifizieren und Dimensionskonflikte über Schichten hinweg zu identifizieren. Das Tool implementiert strukturelle Analysen durch rekursive Modul-Traversierung, Hook-basiertes Tensor-Tracking und eingabegesteuerte Shape-Inferenz. Diese Fähigkeiten ermöglichen die Aggregation von Parameteranzahlen und das Mapping des Datenflusses zwischen aufeinanderfolgenden Operationen.
Determines output dimensions by passing dummy tensors through the network to trigger actual layer computations.
IREE is an MLIR-based compiler toolchain and runtime designed to translate machine learning models from various frameworks into optimized binaries for execution across diverse hardware targets. It provides a unified pipeline to ingest models from PyTorch, TensorFlow, JAX, and ONNX, lowering them into a common intermediate representation for deployment on CPUs, GPUs, and bare-metal embedded systems. The project distinguishes itself through a bytecode virtual machine and a hardware abstraction layer that decouple high-level model logic from specific hardware instruction sets. It supports sophis
Processes tensors with varying dimensions while prioritizing static shapes to maintain high performance.
xtensor is a C++ multidimensional array library for numerical computing that provides N-dimensional containers with an interface mirroring the NumPy API. It utilizes a lazy evaluation expression engine to defer numerical computations until assignment, which minimizes memory allocations and intermediate copies. The library features a foreign memory array adaptor that allows it to wrap external buffers, such as NumPy arrays, to perform numerical operations in-place without duplicating data. It further optimizes performance through lazy broadcasting and a system that manages the lifetime of temp
Expands tensor shapes lazily to execute operations without creating intermediate temporary copies.
ComfyUI-GGUF is a memory optimizer and model loader for ComfyUI that enables the execution of large transformer-based generative models using quantized weights. It provides a system for loading GGUF formatted weights within a node-based diffusion interface to reduce GPU memory consumption. The project includes a quantization tool for converting standard model checkpoints into compressed binary formats and a tensor fixer to restore missing keys and correct architectures in binary model files. These utilities ensure that compressed models remain functional during inference on hardware with limi
Corrects tensor shapes and restores missing keys in model architectures to ensure functional inference.
This project is a deep learning model compiler and parser that converts ONNX models into optimized TensorRT engines. It functions as a bridge that maps standardized ONNX operators to vendor-specific kernels to enable high-performance inference on NVIDIA GPUs. The system operates as a GPU inference optimizer, selecting hardware-specific kernels and tuning memory allocation to maximize throughput. It transforms neural network graphs into serialized binary execution plans to reduce runtime overhead. The toolset covers deep learning model deployment and edge AI performance tuning. It includes ca
Determines optimal memory allocation and tensor dimensions during the build process to support variable input sizes.