7 Repos
Techniques to reduce the number of active parameters during inference to increase token throughput.
Distinct from Model Performance Optimization: Distinct from general performance optimization by focusing specifically on reducing the active parameter set (sparsity) during execution.
Explore 7 awesome GitHub repositories matching artificial intelligence & ml · Model Sparsification. Refine with filters or upvote what's useful.
pykan is a library for implementing Kolmogorov-Arnold Networks, replacing fixed node activation functions with learnable spline functions located on the network edges. It serves as an interpretable AI framework and symbolic regression tool designed to derive transparent mathematical rules from complex data. The project focuses on converting learned numerical functions into human-readable symbolic expressions through library matching and formula conversion. It utilizes additive-compositional topologies and learnable piecewise polynomial segments to approximate non-linear mappings. The framewo
Uses regularization-driven sparsification to force unimportant connections to zero for better interpretability.
PowerInfer is an inference engine and serving framework designed to run large language models on local hardware. It combines a hybrid CPU-GPU offloader, a quantization tool, and a sparse model optimizer to enable the execution of high-parameter models on consumer-grade devices. The system distinguishes itself through neuron-activation-based offloading, using a predictor model to preload frequent neurons into VRAM while keeping rare neurons in system memory. This hybrid execution model balances workloads between the GPU and CPU based on input patterns to optimize memory access and increase tok
Reduces active parameters during execution to increase token throughput and improve processing speed.
MergeKit is a toolkit for combining multiple pre-trained large language models into a single entity using algorithmic blending. It provides a specialized system for parameter interpolation and weight extraction to unify model capabilities. The project distinguishes itself through an evolutionary merge optimizer that tunes parameters based on quantitative evaluation metrics. It also features a mixture of experts orchestrator capable of converting dense models into sparse architectures and a tokenizer alignment tool for transplanting embeddings between different models. The toolkit covers a br
Implements parameter pruning and sign conflict resolution to create sparse model representations.
x-transformers ist eine PyTorch-Bibliothek und ein Research-Toolkit für den Aufbau von Transformer-Architekturen. Es bietet ein modulares Framework für die Implementierung experimenteller Transformer-Forschung, einschließlich einer Suite fortschrittlicher Attention-Mechanismen, Tools für die Modellierung langer Sequenzen und eines Frameworks für Vision-Transformer. Das Projekt zeichnet sich durch den Fokus auf speichereffiziente und performante Komponenten aus, wie etwa Flash-Attention mit Tiled-Kernels und Multi-Query-Attention. Zudem implementiert es spezialisierte Methoden zur Erweiterung von Kontextfenstern, einschließlich Sequence-Recurrence und Rotary-Positional-Embeddings. Die Bibliothek deckt ein breites Spektrum architektonischer Funktionen ab, darunter verschiedene Normalisierungsschemata zur Stabilisierung des Trainings, Gated-Feedforward-Netzwerke und benutzerdefinierte Layer-Topologien wie Macaron-Netzwerke. Sie unterstützt sowohl Encoder- als auch Decoder-Konstruktionen und bietet Tools für die autoregressive Sequenzgenerierung sowie Vision-Language-Aufgaben wie Bildunterschriften.
Implements top-k selection to zero out low-importance attention scores, reducing computational overhead.
Dieses Projekt ist eine PyTorch-Bibliothek für den Aufbau und das Training von Kolmogorov-Arnold-Networks. Es implementiert eine neuronale Netzwerkarchitektur, die feste Aktivierungsfunktionen durch lernbare Spline-basierte Funktionen an den Kanten ersetzt und als Tool für interpretierbares Machine Learning dient. Die Implementierung nutzt reformulierte Matrixoperationen, um den Speicher-Overhead zu reduzieren und die Rechengeschwindigkeit zu erhöhen. Sie verwendet L1-Regularisierung, um Netzwerkgewichte zu spärlich zu machen, was die Transparenz der internen Logik und Entscheidungen des Modells verbessert. Das Framework deckt eine Reihe von Funktionen ab, darunter gitterbasierte Funktionsapproximation, B-Spline-Aktivierungsfunktionen und die Optimierung von Deep-Learning-Modellen. Diese Funktionen sind unter Verwendung nativer PyTorch-Tensoren aufgebaut, um automatische Differenzierung und Hardwarebeschleunigung zu unterstützen.
Includes utilities for model weight sparsification via L1 regularization to improve interpretability.
Model-Optimizer is a deep learning toolkit and framework dedicated to compressing, pruning, quantizing, and optimizing neural network architectures. It provides methodologies covering weight quantization, model distillation, and speculative decoding for efficient text generation, alongside automated neural architecture search for discovering optimal network structures. The library implements post-training quantization pipelines that convert high-precision neural network weights into lower-bit formats using calibration data. Additional optimization techniques include teacher-student knowledge
Transforms pre-trained dense neural network models into sparse variants using magnitude-based thresholding or data-driven calibration without retraining.
PocketFlow is an integrated toolkit for deep learning model compression, distributed training, and mobile format optimization. It provides a system for reducing the size and complexity of neural networks to improve inference efficiency, featuring a dedicated engine for knowledge distillation and a mobile model optimizer. The framework differentiates itself through an automated hyperparameter tuning system that uses reinforcement learning and statistical models to determine optimal compression ratios and layer-wise bit allocation. It also includes a distributed training system that utilizes mu
Implements a dynamic pruning schedule to reduce the number of non-zero weights and decrease inference cost.