10 dépôts
Architectural designs that activate only a subset of parameters per input to improve computational efficiency.
Distinguishing note: Focuses on conditional computation and routing mechanisms, distinct from dense model architectures.
Explore 10 awesome GitHub repositories matching artificial intelligence & ml · Sparse Model Architectures. Refine with filters or upvote what's useful.
This project is a comprehensive framework for the entire lifecycle of transformer-based language models, supporting everything from foundational pretraining to specialized deployment. It provides a modular toolkit for defining neural network architectures, managing data preparation pipelines, and executing training routines across various scales. The framework is designed to handle the full model development process, including supervised fine-tuning, behavioral alignment, and the integration of agentic capabilities. What distinguishes this framework is its focus on efficient training and adva
Computational load is distributed across specialized sub-networks where only a subset of parameters is activated for each input token.
Grok-1 is an open-weights large language model implementation featuring a sparse mixture-of-experts architecture. It is designed for high-performance text generation and natural language processing by activating only a subset of specialized expert layers per token. The model utilizes 8-bit weight quantization to reduce memory overhead and accelerate loading. To manage its high parameter count, the implementation supports activation sharding, which distributes the memory load across multiple hardware devices during execution. The project covers large-scale model inference, including text comp
Employs a sparse architectural design that activates only a subset of parameters per token.
DeepSpeed is a distributed deep learning optimization library and framework designed for the training and inference of massive AI models. It serves as a model parallelism orchestrator and a toolkit for scaling large language models across multiple GPUs and compute nodes. The project distinguishes itself through 3D parallelism orchestration, which combines data, pipeline, and tensor parallelism. It utilizes ZeRO-based memory partitioning to eliminate redundant storage and employs CPU-offload memory management to move weights and optimizer states to system RAM. Additionally, it provides special
Provides specialized routing and support for sparse Mixture-of-Experts architectures to increase model capacity.
PowerInfer is a high-performance local large language model inference engine and sparse inference framework. It provides a runtime for executing models on consumer-grade hardware, utilizing a GPU acceleration backend to optimize tensor operations for graphics processors. The system distinguishes itself through a sparse inference framework that increases generation speed by skipping computations based on activation sparsity in model weights. It includes a GGUF model converter for transforming weights and metadata into a unified binary format, as well as an OpenAI API compatible server for inte
Increases generation speed by identifying and ignoring inactive neurons based on activation sparsity.
Supports Transformer variants, mixture-of-experts, and compression techniques for sparse networks.
gpt-neox is a distributed training system and framework for building large-scale autoregressive language models. It implements the transformer architecture and provides a toolkit for training models with billions of parameters by distributing weights across compute clusters. The framework distinguishes itself through extensive support for distributed model parallelism, including pipeline and sequence parallelism, to overcome single-device memory limits. It further supports sparse model architectures using a mixture of experts system with Sinkhorn-based routing. The project covers a broad ran
Supports sparse model architectures using a mixture of experts system to improve computational efficiency.
Aerosolve est un framework de machine learning conçu pour l'entraînement et le déploiement de modèles interprétables. Il fonctionne comme un outil d'ingénierie des caractéristiques (feature engineering) et un entraîneur de modèles utilisant la modélisation de caractéristiques creuses (sparse feature modeling) pour simplifier le débogage des poids et accélérer l'itération sur les données. Le système inclut un langage de transformation spécifique au domaine pour convertir des données brutes en représentations prêtes pour le modèle. Il offre également des capacités d'analyse de contenu visuel en mappant les images dans des espaces vectoriels denses de haute dimension pour classer et organiser les données par style ou par contenu. Le framework permet un entraînement centré sur l'humain en injectant des croyances a priori et des poids spécifiques dans le processus d'apprentissage. Pour le déploiement, il utilise un runtime d'inférence minimal pour exécuter des prédictions légères et un mécanisme de scoring à contexte partagé pour traiter plusieurs éléments en une seule opération.
Utilizes sparse feature modeling to create interpretable models that simplify weight debugging and iteration.
Engram est un système de récupération de connaissances dynamique et un framework d'augmentation de mémoire pour les grands modèles de langage (LLM). Il fonctionne comme une couche de recherche en mémoire scalable et un composant d'architecture creuse conçu pour fusionner les connaissances statiques du modèle avec des états externes dynamiques afin d'améliorer la véracité et réduire les hallucinations. Le système utilise la récupération conditionnelle en mémoire et l'adressage mémoire différentiable pour mapper les jetons d'entrée vers des indices spécifiques au sein d'un magasin de mémoire associative à grande échelle. Cela permet au modèle d'augmenter ses paramètres totaux disponibles en stockant des poids dans des tables de recherche externes et en n'activant que les segments de connaissances pertinents pour une entrée donnée. Le framework couvre l'optimisation de la sparsité des modèles et l'augmentation scalable, utilisant la récupération clé-valeur et la fusion dynamique de paramètres pour améliorer les performances sur des tâches spécialisées sans nécessiter un réentraînement complet du réseau.
Optimizes memory usage by implementing an architecture that activates only relevant knowledge segments.
Amazon DSSTNE is a machine learning toolkit and sparse tensor network library designed for deep learning models with sparse inputs and outputs. It provides a model-parallel training framework and a GPU-accelerated sparse engine to support memory-intensive networks. The framework is specifically designed for recommendation system training and large-scale sparse learning. It enables the distribution of large weight matrices and embedding tables across multiple GPU devices to handle models that exceed the memory capacity of a single processor. The project covers a broad range of capabilities in
Enables constructing machine learning models using scalable sparse tensor networks to handle large-scale data.
FastDeploy is a high-performance deployment framework for large language models, vision models, and multimodal models. It provides the infrastructure to launch model services that process combined image, video, and text inputs, exposing these capabilities through a standardized, OpenAI-compatible API for chat and text completions. The project distinguishes itself through advanced inference pipeline engineering and GPU optimization. It employs speculative decoding, tensor parallelism, and a disaggregated execution model that separates prefill and decode phases across different hardware resourc
Uses sparse attention mechanisms to process key-value blocks selectively and handle long-sequence inputs.