10 个仓库
Architectural designs that activate only a subset of parameters per input to improve computational efficiency.
Distinguishing note: Focuses on conditional computation and routing mechanisms, distinct from dense model architectures.
Explore 10 awesome GitHub repositories matching artificial intelligence & ml · Sparse Model Architectures. Refine with filters or upvote what's useful.
This project is a comprehensive framework for the entire lifecycle of transformer-based language models, supporting everything from foundational pretraining to specialized deployment. It provides a modular toolkit for defining neural network architectures, managing data preparation pipelines, and executing training routines across various scales. The framework is designed to handle the full model development process, including supervised fine-tuning, behavioral alignment, and the integration of agentic capabilities. What distinguishes this framework is its focus on efficient training and adva
Computational load is distributed across specialized sub-networks where only a subset of parameters is activated for each input token.
Grok-1 is an open-weights large language model implementation featuring a sparse mixture-of-experts architecture. It is designed for high-performance text generation and natural language processing by activating only a subset of specialized expert layers per token. The model utilizes 8-bit weight quantization to reduce memory overhead and accelerate loading. To manage its high parameter count, the implementation supports activation sharding, which distributes the memory load across multiple hardware devices during execution. The project covers large-scale model inference, including text comp
Employs a sparse architectural design that activates only a subset of parameters per token.
DeepSpeed is a distributed deep learning optimization library and framework designed for the training and inference of massive AI models. It serves as a model parallelism orchestrator and a toolkit for scaling large language models across multiple GPUs and compute nodes. The project distinguishes itself through 3D parallelism orchestration, which combines data, pipeline, and tensor parallelism. It utilizes ZeRO-based memory partitioning to eliminate redundant storage and employs CPU-offload memory management to move weights and optimizer states to system RAM. Additionally, it provides special
Provides specialized routing and support for sparse Mixture-of-Experts architectures to increase model capacity.
PowerInfer is a high-performance local large language model inference engine and sparse inference framework. It provides a runtime for executing models on consumer-grade hardware, utilizing a GPU acceleration backend to optimize tensor operations for graphics processors. The system distinguishes itself through a sparse inference framework that increases generation speed by skipping computations based on activation sparsity in model weights. It includes a GGUF model converter for transforming weights and metadata into a unified binary format, as well as an OpenAI API compatible server for inte
Increases generation speed by identifying and ignoring inactive neurons based on activation sparsity.
Supports Transformer variants, mixture-of-experts, and compression techniques for sparse networks.
gpt-neox is a distributed training system and framework for building large-scale autoregressive language models. It implements the transformer architecture and provides a toolkit for training models with billions of parameters by distributing weights across compute clusters. The framework distinguishes itself through extensive support for distributed model parallelism, including pipeline and sequence parallelism, to overcome single-device memory limits. It further supports sparse model architectures using a mixture of experts system with Sinkhorn-based routing. The project covers a broad ran
Supports sparse model architectures using a mixture of experts system to improve computational efficiency.
Aerosolve is a machine learning framework designed for training and deploying interpretable models. It functions as a feature engineering tool and a model trainer that utilizes sparse feature modeling to simplify weight debugging and accelerate data iteration. The system includes a specialized domain-specific transformation language for converting raw data families into model-ready representations. It also provides capabilities for visual content analysis by mapping images into dense high-dimensional vector spaces to rank and organize data by style or content. The framework allows for human-
Utilizes sparse feature modeling to create interpretable models that simplify weight debugging and iteration.
Engram 是一个用于大语言模型的动态知识检索系统和记忆增强框架。它作为一个可扩展的记忆查找层和稀疏架构组件,旨在将静态模型知识与动态外部状态融合,以提高事实准确性并减少幻觉。 该系统利用条件记忆检索和可微记忆寻址,将输入 Token 映射到大规模关联记忆存储中的特定索引。这允许模型通过将权重存储在外部查找表中,并仅为给定输入激活相关的知识片段,从而增加其总可用参数量。 该框架涵盖了模型稀疏性优化和可扩展增强,使用键值检索和动态参数融合来提升特定任务的性能,而无需对网络进行全面重训练。
Optimizes memory usage by implementing an architecture that activates only relevant knowledge segments.
Amazon DSSTNE 是一个机器学习工具包和稀疏张量网络库,专为具有稀疏输入和输出的深度学习模型而设计。它提供了一个模型并行训练框架和一个 GPU 加速的稀疏引擎,以支持内存密集型网络。 该框架专门为推荐系统训练和大规模稀疏学习而设计。它实现了将大型权重矩阵和嵌入表分布在多个 GPU 设备上,以处理超过单个处理器内存容量的模型。 该项目涵盖了广泛的能力,包括分布式 GPU 计算、稀疏数据集处理以及可扩展稀疏张量网络的构建。这些实用程序允许在 GPU 集群上执行高性能机器学习操作和模型扩展。
Enables constructing machine learning models using scalable sparse tensor networks to handle large-scale data.
FastDeploy is a high-performance deployment framework for large language models, vision models, and multimodal models. It provides the infrastructure to launch model services that process combined image, video, and text inputs, exposing these capabilities through a standardized, OpenAI-compatible API for chat and text completions. The project distinguishes itself through advanced inference pipeline engineering and GPU optimization. It employs speculative decoding, tensor parallelism, and a disaggregated execution model that separates prefill and decode phases across different hardware resourc
Uses sparse attention mechanisms to process key-value blocks selectively and handle long-sequence inputs.