awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

10 repositorios

Awesome GitHub RepositoriesSparse Model Architectures

Architectural designs that activate only a subset of parameters per input to improve computational efficiency.

Distinguishing note: Focuses on conditional computation and routing mechanisms, distinct from dense model architectures.

Explore 10 awesome GitHub repositories matching artificial intelligence & ml · Sparse Model Architectures. Refine with filters or upvote what's useful.

Awesome Sparse Model Architectures GitHub Repositories

Encuentra los mejores repositorios con IA.Buscaremos los repositorios que mejor coincidan usando IA.
  • jingyaogong/minimindAvatar de jingyaogong

    jingyaogong/minimind

    51,834Ver en GitHub↗

    This project is a comprehensive framework for the entire lifecycle of transformer-based language models, supporting everything from foundational pretraining to specialized deployment. It provides a modular toolkit for defining neural network architectures, managing data preparation pipelines, and executing training routines across various scales. The framework is designed to handle the full model development process, including supervised fine-tuning, behavioral alignment, and the integration of agentic capabilities. What distinguishes this framework is its focus on efficient training and adva

    Computational load is distributed across specialized sub-networks where only a subset of parameters is activated for each input token.

    Pythonartificial-intelligencelarge-language-model
    Ver en GitHub↗51,834
  • xai-org/grok-1Avatar de xai-org

    xai-org/grok-1

    51,690Ver en GitHub↗

    Grok-1 is an open-weights large language model implementation featuring a sparse mixture-of-experts architecture. It is designed for high-performance text generation and natural language processing by activating only a subset of specialized expert layers per token. The model utilizes 8-bit weight quantization to reduce memory overhead and accelerate loading. To manage its high parameter count, the implementation supports activation sharding, which distributes the memory load across multiple hardware devices during execution. The project covers large-scale model inference, including text comp

    Employs a sparse architectural design that activates only a subset of parameters per token.

    Python
    Ver en GitHub↗51,690
  • microsoft/deepspeedAvatar de microsoft

    microsoft/DeepSpeed

    42,533Ver en GitHub↗

    DeepSpeed is a distributed deep learning optimization library and framework designed for the training and inference of massive AI models. It serves as a model parallelism orchestrator and a toolkit for scaling large language models across multiple GPUs and compute nodes. The project distinguishes itself through 3D parallelism orchestration, which combines data, pipeline, and tensor parallelism. It utilizes ZeRO-based memory partitioning to eliminate redundant storage and employs CPU-offload memory management to move weights and optimizer states to system RAM. Additionally, it provides special

    Provides specialized routing and support for sparse Mixture-of-Experts architectures to increase model capacity.

    Python
    Ver en GitHub↗42,533
  • tiiny-ai/powerinferAvatar de Tiiny-AI

    Tiiny-AI/PowerInfer

    8,714Ver en GitHub↗

    PowerInfer is a high-performance local large language model inference engine and sparse inference framework. It provides a runtime for executing models on consumer-grade hardware, utilizing a GPU acceleration backend to optimize tensor operations for graphics processors. The system distinguishes itself through a sparse inference framework that increases generation speed by skipping computations based on activation sparsity in model weights. It includes a GGUF model converter for transforming weights and metadata into a unified binary format, as well as an OpenAI API compatible server for inte

    Increases generation speed by identifying and ignoring inactive neurons based on activation sparsity.

    C++large-language-modelsllamallm
    Ver en GitHub↗8,714
  • infrasys-ai/aiinfraAvatar de Infrasys-AI

    Infrasys-AI/AIInfra

    7,414Ver en GitHub↗

    Supports Transformer variants, mixture-of-experts, and compression techniques for sparse networks.

    Jupyter Notebookaiinfraaisystem
    Ver en GitHub↗7,414
  • eleutherai/gpt-neoxAvatar de EleutherAI

    EleutherAI/gpt-neox

    7,392Ver en GitHub↗

    gpt-neox is a distributed training system and framework for building large-scale autoregressive language models. It implements the transformer architecture and provides a toolkit for training models with billions of parameters by distributing weights across compute clusters. The framework distinguishes itself through extensive support for distributed model parallelism, including pipeline and sequence parallelism, to overcome single-device memory limits. It further supports sparse model architectures using a mixture of experts system with Sinkhorn-based routing. The project covers a broad ran

    Supports sparse model architectures using a mixture of experts system to improve computational efficiency.

    Pythondeepspeed-librarygpt-3language-model
    Ver en GitHub↗7,392
  • airbnb/aerosolveAvatar de airbnb

    airbnb/aerosolve

    4,804Ver en GitHub↗

    Aerosolve es un framework de machine learning diseñado para entrenar y desplegar modelos interpretables. Funciona como una herramienta de ingeniería de características y un entrenador de modelos que utiliza modelado de características dispersas para simplificar la depuración de pesos y acelerar la iteración de datos. El sistema incluye un lenguaje de transformación específico del dominio para convertir familias de datos crudos en representaciones listas para el modelo. También proporciona capacidades para el análisis de contenido visual mediante el mapeo de imágenes en espacios vectoriales densos de alta dimensión para clasificar y organizar datos por estilo o contenido. El framework permite un entrenamiento centrado en el humano al inyectar creencias previas y pesos específicos en el proceso de aprendizaje del modelo. Para el despliegue, utiliza un runtime de inferencia mínimo para ejecutar predicciones ligeras y un mecanismo de puntuación de contexto compartido para procesar múltiples elementos en una sola operación.

    Utilizes sparse feature modeling to create interpretable models that simplify weight debugging and iteration.

    Scala
    Ver en GitHub↗4,804
  • deepseek-ai/engramAvatar de deepseek-ai

    deepseek-ai/Engram

    4,462Ver en GitHub↗

    Engram es un sistema dinámico de recuperación de conocimiento y framework de aumento de memoria para modelos de lenguaje grandes (LLM). Funciona como una capa de búsqueda de memoria escalable y un componente de arquitectura dispersa diseñado para fusionar el conocimiento estático del modelo con estados externos dinámicos para mejorar la veracidad y reducir las alucinaciones. El sistema utiliza recuperación de memoria condicional y direccionamiento de memoria diferenciable para mapear tokens de entrada a índices específicos dentro de un almacén de memoria asociativa a gran escala. Esto permite al modelo aumentar sus parámetros totales disponibles almacenando pesos en tablas de búsqueda externas y activando solo los segmentos de conocimiento relevantes para una entrada dada. El framework cubre la optimización de dispersión del modelo y el aumento escalable, utilizando recuperación clave-valor y fusión dinámica de parámetros para mejorar el rendimiento en tareas especializadas sin requerir un reentrenamiento completo de la red.

    Optimizes memory usage by implementing an architecture that activates only relevant knowledge segments.

    Python
    Ver en GitHub↗4,462
  • amznlabs/amazon-dsstneAvatar de amznlabs

    amznlabs/amazon-dsstne

    4,395Ver en GitHub↗

    Amazon DSSTNE es un kit de herramientas de machine learning y librería de redes de tensores dispersos diseñada para modelos de deep learning con entradas y salidas dispersas. Proporciona un framework de entrenamiento paralelo al modelo y un motor disperso acelerado por GPU para soportar redes intensivas en memoria. El framework está diseñado específicamente para el entrenamiento de sistemas de recomendación y aprendizaje disperso a gran escala. Permite la distribución de grandes matrices de pesos y tablas de embedding a través de múltiples dispositivos GPU para manejar modelos que exceden la capacidad de memoria de un solo procesador. El proyecto cubre una amplia gama de capacidades, incluyendo computación distribuida en GPU, procesamiento de datasets dispersos y la construcción de redes de tensores dispersos escalables. Estas utilidades permiten la ejecución de operaciones de machine learning de alto rendimiento y el escalado de modelos a través de clústeres de GPU.

    Enables constructing machine learning models using scalable sparse tensor networks to handle large-scale data.

    C++
    Ver en GitHub↗4,395
  • paddlepaddle/fastdeployAvatar de PaddlePaddle

    PaddlePaddle/FastDeploy

    3,700Ver en GitHub↗

    FastDeploy is a high-performance deployment framework for large language models, vision models, and multimodal models. It provides the infrastructure to launch model services that process combined image, video, and text inputs, exposing these capabilities through a standardized, OpenAI-compatible API for chat and text completions. The project distinguishes itself through advanced inference pipeline engineering and GPU optimization. It employs speculative decoding, tensor parallelism, and a disaggregated execution model that separates prefill and decode phases across different hardware resourc

    Uses sparse attention mechanisms to process key-value blocks selectively and handle long-sequence inputs.

    Pythonernieernie-45ernie-45-vl
    Ver en GitHub↗3,700
  1. Home
  2. Artificial Intelligence & ML
  3. Sparse Model Architectures

Explorar subetiquetas

  • Sparse Inference FrameworksExecution systems that optimize inference by skipping computations based on activation sparsity in model weights. **Distinct from Sparse Model Architectures:** Distinct from Sparse Model Architectures: focuses on the execution runtime and skipping computations rather than the neural network design.
  • Sparse Tensor Network ModelingThe process of designing model architectures based on sparse tensor network topologies. **Distinct from Sparse Model Architectures:** Focuses on the network topology of sparse tensors rather than general conditional computation or routing.