awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

13 repositorios

Awesome GitHub RepositoriesLow-Bit Weight Quantization

Converting model weights to low-bit precision formats like INT4 and FP8.

Distinct from Intel Hardware Acceleration: Focuses on weight quantization for AI models specifically, rather than general GPU video decoding acceleration.

Explore 13 awesome GitHub repositories matching devops & infrastructure · Low-Bit Weight Quantization. Refine with filters or upvote what's useful.

Awesome Low-Bit Weight Quantization GitHub Repositories

Encuentra los mejores repositorios con IA.Buscaremos los repositorios que mejor coincidan usando IA.
  • fminference/flexllmgenAvatar de FMInference

    FMInference/FlexLLMGen

    9,362Ver en GitHub↗

    FlexLLMGen is an inference engine and runtime designed to run large language models on a single GPU by combining weight compression with tensor offloading. It reduces model weight memory usage by approximately 70% through 4-bit quantization, and stores model parameters, attention cache, and hidden states across GPU, CPU, and disk to fit models larger than available GPU memory. The project distinguishes itself through a throughput-oriented batching approach that processes multiple generation requests together in large batches to maximize throughput on a single GPU. It also supports distributed

    Reduces model weight memory by approximately 70% using 4-bit quantization with minimal accuracy loss.

    Pythondeep-learninggpt-3high-throughput
    Ver en GitHub↗9,362
  • intel-analytics/bigdlAvatar de intel-analytics

    intel-analytics/BigDL

    8,845Ver en GitHub↗

    BigDL es un framework de aceleración de PyTorch y motor de inferencia distribuida diseñado para grandes modelos de lenguaje. Proporciona un kit de herramientas para ejecutar modelos en hardware Intel, integrando herramientas de cuantización y librerías para el ajuste fino eficiente en parámetros. El proyecto se distingue por el uso de paralelismo de pipeline para distribuir cargas de trabajo de modelos a través de múltiples aceleradores de hardware. Utiliza cuantización de enteros de bajo bit y decodificación especulativa para reducir la huella de memoria y disminuir la latencia de generación de texto. El sistema cubre amplias capacidades en optimización de modelos, incluyendo compresión de pesos y carga de modelos cuantizados. También admite rutinas de entrenamiento aceleradas por hardware para adaptar modelos preentrenados a tareas específicas.

    Compresses LLM weights into low-bit precision formats to reduce memory usage and increase execution speed.

    Python
    Ver en GitHub↗8,845
  • intel/ipex-llmAvatar de intel

    intel/ipex-llm

    8,836Ver en GitHub↗

    Intel XPU LLM Acceleration Library is a toolkit designed to accelerate large language model inference and finetuning on Intel CPUs, GPUs, and NPUs. It provides a distributed inference engine for scaling models across multiple accelerators, a multimodal model runtime for vision and speech tasks, and a low-bit model quantization tool for converting weights into INT4, FP8, and GGUF formats. The project features a parameter-efficient finetuning framework that enables model adaptation using QLoRA and DPO on Intel hardware. It distinguishes itself by providing specialized optimizations for Intel XP

    Converts model weights to low-bit precision formats like INT4 and FP8 to maximize performance on Intel hardware.

    Python
    Ver en GitHub↗8,836
  • 01-ai/yiAvatar de 01-ai

    01-ai/Yi

    7,822Ver en GitHub↗

    Yi is a bilingual language model and foundation model designed for natural language processing, reasoning, and reading comprehension in both English and Chinese. It is built as a transformer-based architecture capable of general purpose text generation and conversational tasks. The model is distinguished by its ability to function as a long context system, processing and analyzing extended input sequences up to 200k tokens. It also supports quantized versions that use low-bit precision to reduce memory footprints, enabling execution on consumer-grade hardware. The project covers a broad rang

    Provides low-bit weight quantization to reduce memory footprint for execution on consumer-grade hardware.

    Jupyter Notebooklarge-language-models
    Ver en GitHub↗7,822
  • ericlbuehler/mistral.rsAvatar de EricLBuehler

    EricLBuehler/mistral.rs

    6,597Ver en GitHub↗

    mistral.rs is an inference engine for large language models that runs locally and exposes models behind OpenAI and Anthropic-compatible APIs. It serves as a multi-model serving platform, capable of loading several models in a single server process with per-request routing and on-demand loading and unloading. The engine supports multimodal inference, processing text alongside images, video, audio, and speech inputs, and includes a quantized model deployment runtime that reduces memory use and speeds up inference on consumer hardware. The project distinguishes itself through an agentic tool exe

    Uses per-column weight data from calibration text to allocate more precision to high-impact weights during quantization.

    Rustllmrustuqff
    Ver en GitHub↗6,597
  • lightning-ai/lit-llamaAvatar de Lightning-AI

    Lightning-AI/lit-llama

    6,081Ver en GitHub↗

    Lit-llama es un framework de implementación basado en PyTorch para el modelo de lenguaje LLaMA, que proporciona un sistema para pre-entrenamiento, ajuste fino e inferencia de alto rendimiento. Incluye un pipeline de pre-entrenamiento para crear modelos de lenguaje fundamentales desde cero y herramientas para ejecutar pesos pre-entrenados para generar texto natural y predecir secuencias. El proyecto proporciona toolkits especializados para el ajuste fino eficiente en parámetros utilizando adaptación de bajo rango (LoRA) y adaptadores ligeros. También incluye una librería de cuantización que reduce las huellas de memoria del modelo a través de precisión de cuatro y ocho bits para permitir la ejecución en hardware con recursos limitados. El framework incorpora un diseño de transformador simplificado y emplea flash attention para optimizar la memoria y la velocidad. Además, gestiona datasets a gran escala a través de formatos de datos en streaming para evitar cargar corpora completos en la memoria del sistema.

    Ships a quantization library that reduces memory footprints via GPTQ-based 4-bit and 8-bit precision.

    Python
    Ver en GitHub↗6,081
  • ace-step/ace-step-1.5Avatar de ace-step

    ace-step/ACE-Step-1.5

    6,002Ver en GitHub↗

    ACE Step 1.5 is a local text-to-music generation and audio editing system that runs on consumer hardware. It transforms plain-language descriptions into full-length songs with lyrics, and can edit existing audio through cover generation, vocal removal, track separation, and selective repainting. The system supports multilingual prompts and lyrics in over 50 languages, and provides precise control over musical structure including duration, BPM, key, and time signature. The project distinguishes itself through a dual-stream diffusion architecture that processes separate latent streams for vocal

    Suno generates complete songs in under ten seconds on a standard consumer GPU while using less than four gigabytes of video memory.

    Python
    Ver en GitHub↗6,002
  • baichuan-inc/baichuan-7bAvatar de baichuan-inc

    baichuan-inc/Baichuan-7B

    5,654Ver en GitHub↗

    Baichuan-7B is an open-source 7 billion parameter bilingual Transformer model designed for text generation and few-shot learning across Chinese and English. It is built on a large Transformer architecture trained on a bilingual corpus, enabling it to produce coherent text in both languages from a single model. The model incorporates several optimization techniques that distinguish it from standard large language models. It uses rotary position embeddings that can extrapolate to longer sequences than seen during training, allowing context extension beyond the original 4096-token training lengt

    Reduces model memory by approximately 70% using 4-bit weight quantization with minimal accuracy loss.

    Pythonartificial-intelligencecevalchatgpt
    Ver en GitHub↗5,654
  • panqiwei/autogptqAvatar de PanQiWei

    PanQiWei/AutoGPTQ

    5,073Ver en GitHub↗

    AutoGPTQ es un framework de compresión de modelos diseñado para reducir la huella de memoria y aumentar la velocidad de inferencia de modelos de lenguaje grandes (LLM). Utiliza el algoritmo GPTQ para comprimir los pesos del modelo, permitiendo que estos modelos se ejecuten en hardware con VRAM limitada. El kit de herramientas proporciona una tubería de cuantización de arquitectura que admite la integración de clases de modelos personalizados para diversas arquitecturas de redes neuronales. Incluye un motor de inferencia de precisión mixta con kernels optimizados para acelerar la multiplicación de matrices durante el despliegue. El framework cubre el flujo de trabajo completo de compresión de pesos, desde la calibración y cuantización hasta la evaluación de precisión posterior. Estas herramientas miden la pérdida de rendimiento comparando la salida de los modelos cuantizados contra los pesos originales en tareas de referencia.

    Implements the GPTQ algorithm for post-training weight quantization to reduce model size.

    Python
    Ver en GitHub↗5,073
  • autogptq/autogptqAvatar de AutoGPTQ

    AutoGPTQ/AutoGPTQ

    5,070Ver en GitHub↗

    AutoGPTQ es un kit de herramientas de compresión de modelos y un framework de cuantización post-entrenamiento diseñado para reducir la huella de memoria de modelos de lenguaje grandes. Utiliza el algoritmo GPTQ para comprimir los pesos de las redes neuronales, reduciendo los requisitos de hardware y el uso de VRAM. El proyecto sirve como un acelerador de inferencia al proporcionar kernels optimizados que aumentan la velocidad de generación de tokens. Cuenta con extensibilidad de arquitectura de modelo, permitiendo que las capacidades de cuantización se añadan a nuevas estructuras de modelos mediante patrones configurables. El framework cubre una tubería de cuantización integral, incluyendo compresión de pesos por capa, estimación de escala basada en calibración y mapeo de memoria específico de precisión. También incluye sistemas para la evaluación del rendimiento del modelo para medir el impacto de la cuantización en la precisión en tareas de lenguaje y resumen.

    Implements the GPTQ algorithm for high-efficiency post-training weight quantization of large language models.

    Python
    Ver en GitHub↗5,070
  • sakurallm/sakurallmAvatar de SakuraLLM

    SakuraLLM/SakuraLLM

    4,618Ver en GitHub↗

    SakuraLLM is a multi-format document translation system that hosts large language models for translating Japanese text into other languages. It functions as an inference server that exposes translation models through an OpenAI-compatible API, allowing any tool supporting the OpenAI client format to send translation requests. The system is designed as a glossary-aware translation engine that applies user-defined term dictionaries to ensure consistent translation of proper nouns and names across outputs. The project distinguishes itself by supporting multiple high-performance inference backends

    Runs the translation model on NVIDIA and AMD GPUs with CPU-GPU hybrid inference for lower-memory setups.

    Python
    Ver en GitHub↗4,618
  • baichuan-inc/baichuan2Avatar de baichuan-inc

    baichuan-inc/Baichuan2

    4,098Ver en GitHub↗

    Baichuan2 es una colección de modelos de lenguaje grandes pre-entrenados, incluyendo variantes base y de chat, diseñados para la generación de lenguaje natural y la IA conversacional multi-turno. Proporciona un motor de inferencia y un framework de ajuste fino (fine-tuning) para adaptar estos modelos a datasets personalizados y dominios especializados. El proyecto cuenta con un kit de herramientas de cuantización y un motor de inferencia que permiten la ejecución del modelo en hardware diverso, incluyendo procesadores gráficos, procesadores centrales y aceleradores especializados. Estas herramientas admiten la cuantización de pesos de bajo bit para reducir el uso de memoria y aumentar la velocidad de inferencia en hardware limitado. El sistema cubre una amplia gama de capacidades, incluyendo entrenamiento distribuido en múltiples máquinas, ajuste fino eficiente en parámetros y alineación supervisada para la interacción humana. También incluye utilidades para la conversión de versiones de modelos y proporciona interfaces conversacionales a través de herramientas de línea de comandos o demostraciones basadas en web.

    Implements weight quantization to four or eight bits to reduce memory overhead and increase inference speed.

    Pythonartificial-intelligencebenchmarkceval
    Ver en GitHub↗4,098
  • nunchaku-ai/nunchakuAvatar de nunchaku-ai

    nunchaku-ai/nunchaku

    3,883Ver en GitHub↗

    Nunchaku is a 4-bit model quantization library and diffusion model inference engine designed to run large-scale neural networks on consumer GPUs. It functions as a GPU-accelerated optimizer that reduces VRAM usage and increases inference speed through weight compression and memory management. The project utilizes low-rank weight decomposition and SVD weight quantization to compress models to four-bit precision while maintaining visual fidelity. It employs kernel-level operator fusion to minimize data movement and hardware-aware precision mapping to adjust numerical precision based on the unde

    Provides a toolkit for compressing neural network weights into four-bit precision to reduce VRAM usage.

    Pythoncomfyuidiffusion-modelsflux
    Ver en GitHub↗3,883
  1. Home
  2. DevOps & Infrastructure
  3. Intel Hardware Acceleration
  4. Low-Bit Weight Quantization

Explorar subetiquetas

  • 4-Bit Quantization Tools2 sub-etiquetasTools that reduce model weight memory by approximately 70% using 4-bit quantization with minimal accuracy loss. **Distinct from Low-Bit Weight Quantization:** Distinct from Low-Bit Weight Quantization: focuses specifically on 4-bit precision rather than general low-bit formats like INT4 and FP8.
  • Consumer GPU Optimizations1 sub-etiquetaSpecialized configurations to allow large models to run on consumer-grade graphics cards. **Distinct from Low-Bit Weight Quantization:** Focuses on the execution target (consumer hardware) rather than the quantization process itself.
  • Importance Matrix CalibrationTechniques that use per-column weight data from calibration text to allocate more precision to high-impact weights during quantization. **Distinct from Low-Bit Weight Quantization:** Distinct from Low-Bit Weight Quantization: focuses on the calibration method for allocating precision, not the general technique of reducing bit width.