awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

13 مستودعات

Awesome GitHub RepositoriesLow-Bit Weight Quantization

Converting model weights to low-bit precision formats like INT4 and FP8.

Distinct from Intel Hardware Acceleration: Focuses on weight quantization for AI models specifically, rather than general GPU video decoding acceleration.

Explore 13 awesome GitHub repositories matching devops & infrastructure · Low-Bit Weight Quantization. Refine with filters or upvote what's useful.

Awesome Low-Bit Weight Quantization GitHub Repositories

اعثر على أفضل المستودعات باستخدام الذكاء الاصطناعي.سنبحث عن أفضل المستودعات المطابقة باستخدام الذكاء الاصطناعي.
  • fminference/flexllmgenالصورة الرمزية لـ FMInference

    FMInference/FlexLLMGen

    9,362عرض على GitHub↗

    FlexLLMGen is an inference engine and runtime designed to run large language models on a single GPU by combining weight compression with tensor offloading. It reduces model weight memory usage by approximately 70% through 4-bit quantization, and stores model parameters, attention cache, and hidden states across GPU, CPU, and disk to fit models larger than available GPU memory. The project distinguishes itself through a throughput-oriented batching approach that processes multiple generation requests together in large batches to maximize throughput on a single GPU. It also supports distributed

    Reduces model weight memory by approximately 70% using 4-bit quantization with minimal accuracy loss.

    Pythondeep-learninggpt-3high-throughput
    عرض على GitHub↗9,362
  • intel-analytics/bigdlالصورة الرمزية لـ intel-analytics

    intel-analytics/BigDL

    8,845عرض على GitHub↗

    BigDL is a PyTorch acceleration framework and distributed inference engine designed for large language models. It provides a toolkit for running models on Intel hardware, integrating quantization tools and libraries for parameter-efficient fine-tuning. The project distinguishes itself through the use of pipeline parallelism to distribute model workloads across multiple hardware accelerators. It utilizes low-bit integer quantization and speculative decoding to reduce memory footprints and decrease text generation latency. The system covers broad capabilities in model optimization, including w

    Compresses LLM weights into low-bit precision formats to reduce memory usage and increase execution speed.

    Python
    عرض على GitHub↗8,845
  • intel/ipex-llmالصورة الرمزية لـ intel

    intel/ipex-llm

    8,836عرض على GitHub↗

    Intel XPU LLM Acceleration Library is a toolkit designed to accelerate large language model inference and finetuning on Intel CPUs, GPUs, and NPUs. It provides a distributed inference engine for scaling models across multiple accelerators, a multimodal model runtime for vision and speech tasks, and a low-bit model quantization tool for converting weights into INT4, FP8, and GGUF formats. The project features a parameter-efficient finetuning framework that enables model adaptation using QLoRA and DPO on Intel hardware. It distinguishes itself by providing specialized optimizations for Intel XP

    Converts model weights to low-bit precision formats like INT4 and FP8 to maximize performance on Intel hardware.

    Python
    عرض على GitHub↗8,836
  • 01-ai/yiالصورة الرمزية لـ 01-ai

    01-ai/Yi

    7,822عرض على GitHub↗

    Yi is a bilingual language model and foundation model designed for natural language processing, reasoning, and reading comprehension in both English and Chinese. It is built as a transformer-based architecture capable of general purpose text generation and conversational tasks. The model is distinguished by its ability to function as a long context system, processing and analyzing extended input sequences up to 200k tokens. It also supports quantized versions that use low-bit precision to reduce memory footprints, enabling execution on consumer-grade hardware. The project covers a broad rang

    Provides low-bit weight quantization to reduce memory footprint for execution on consumer-grade hardware.

    Jupyter Notebooklarge-language-models
    عرض على GitHub↗7,822
  • ericlbuehler/mistral.rsالصورة الرمزية لـ EricLBuehler

    EricLBuehler/mistral.rs

    6,597عرض على GitHub↗

    mistral.rs is an inference engine for large language models that runs locally and exposes models behind OpenAI and Anthropic-compatible APIs. It serves as a multi-model serving platform, capable of loading several models in a single server process with per-request routing and on-demand loading and unloading. The engine supports multimodal inference, processing text alongside images, video, audio, and speech inputs, and includes a quantized model deployment runtime that reduces memory use and speeds up inference on consumer hardware. The project distinguishes itself through an agentic tool exe

    Uses per-column weight data from calibration text to allocate more precision to high-impact weights during quantization.

    Rustllmrustuqff
    عرض على GitHub↗6,597
  • lightning-ai/lit-llamaالصورة الرمزية لـ Lightning-AI

    Lightning-AI/lit-llama

    6,081عرض على GitHub↗

    Lit-llama هو إطار عمل تنفيذ يعتمد على PyTorch لنموذج اللغة LLaMA، ويوفر نظاماً للتدريب المسبق، والضبط الدقيق، والاستدلال عالي الأداء. يتضمن خط أنابيب تدريب مسبق لإنشاء نماذج لغوية أساسية من الصفر وأدوات لتشغيل الأوزان المدربة مسبقاً لتوليد نص طبيعي والتنبؤ بالتسلسلات. يوفر المشروع مجموعات أدوات متخصصة للضبط الدقيق الفعال للمعلمات باستخدام التكيف منخفض الرتبة (LoRA) والمحولات خفيفة الوزن. كما يتضمن مكتبة تكميم (quantization) تقلل من بصمات ذاكرة النموذج من خلال دقة 4 بت و8 بت لتمكين التنفيذ على الأجهزة ذات الموارد المحدودة. يدمج إطار العمل تصميم محول مبسط ويوظف انتباه الفلاش (flash attention) لتحسين الذاكرة والسرعة. كما يدير مجموعات بيانات واسعة النطاق من خلال تنسيقات بيانات البث لتجنب تحميل مجموعات النصوص الكاملة في ذاكرة النظام.

    Ships a quantization library that reduces memory footprints via GPTQ-based 4-bit and 8-bit precision.

    Python
    عرض على GitHub↗6,081
  • ace-step/ace-step-1.5الصورة الرمزية لـ ace-step

    ace-step/ACE-Step-1.5

    6,002عرض على GitHub↗

    ACE Step 1.5 is a local text-to-music generation and audio editing system that runs on consumer hardware. It transforms plain-language descriptions into full-length songs with lyrics, and can edit existing audio through cover generation, vocal removal, track separation, and selective repainting. The system supports multilingual prompts and lyrics in over 50 languages, and provides precise control over musical structure including duration, BPM, key, and time signature. The project distinguishes itself through a dual-stream diffusion architecture that processes separate latent streams for vocal

    Suno generates complete songs in under ten seconds on a standard consumer GPU while using less than four gigabytes of video memory.

    Python
    عرض على GitHub↗6,002
  • baichuan-inc/baichuan-7bالصورة الرمزية لـ baichuan-inc

    baichuan-inc/Baichuan-7B

    5,654عرض على GitHub↗

    Baichuan-7B is an open-source 7 billion parameter bilingual Transformer model designed for text generation and few-shot learning across Chinese and English. It is built on a large Transformer architecture trained on a bilingual corpus, enabling it to produce coherent text in both languages from a single model. The model incorporates several optimization techniques that distinguish it from standard large language models. It uses rotary position embeddings that can extrapolate to longer sequences than seen during training, allowing context extension beyond the original 4096-token training lengt

    Reduces model memory by approximately 70% using 4-bit weight quantization with minimal accuracy loss.

    Pythonartificial-intelligencecevalchatgpt
    عرض على GitHub↗5,654
  • panqiwei/autogptqالصورة الرمزية لـ PanQiWei

    PanQiWei/AutoGPTQ

    5,073عرض على GitHub↗

    AutoGPTQ هو إطار عمل لضغط النماذج مصمم لتقليل بصمة الذاكرة وزيادة سرعة الاستنتاج لنماذج اللغات الكبيرة. يستخدم خوارزمية GPTQ لضغط أوزان النموذج، مما يسمح لهذه النماذج بالعمل على أجهزة ذات ذاكرة فيديو (VRAM) محدودة. توفر مجموعة الأدوات خط أنابيب لتكميم البنية (Architecture quantization) يدعم دمج فئات النماذج المخصصة لبنيات الشبكات العصبية المختلفة. يتضمن محرك استنتاج بدقة مختلطة مع نواة محسنة لتسريع ضرب المصفوفات أثناء النشر. يغطي إطار العمل سير عمل ضغط الأوزان بالكامل، من المعايرة والتكميم إلى تقييم الدقة اللاحق. تقيس هذه الأدوات فقدان الأداء من خلال مقارنة مخرجات النماذج المكممة مقابل الأوزان الأصلية في مهام القياس المعياري.

    Implements the GPTQ algorithm for post-training weight quantization to reduce model size.

    Python
    عرض على GitHub↗5,073
  • autogptq/autogptqالصورة الرمزية لـ AutoGPTQ

    AutoGPTQ/AutoGPTQ

    5,070عرض على GitHub↗

    AutoGPTQ هو مجموعة أدوات لضغط النماذج وإطار عمل للتكميم بعد التدريب مصمم لتقليل بصمة الذاكرة لنماذج اللغات الكبيرة. يستخدم خوارزمية GPTQ لضغط أوزان الشبكة العصبية، مما يقلل من متطلبات الأجهزة ويقلل من استخدام ذاكرة الفيديو (VRAM). يعمل المشروع كمسرع للاستنتاج من خلال توفير نواة محسنة تزيد من سرعة توليد الرموز (Tokens). يتميز بقابلية توسيع بنية النموذج، مما يسمح بإضافة قدرات التكميم إلى هياكل النماذج الجديدة من خلال أنماط قابلة للتكوين. يغطي إطار العمل خط أنابيب تكميم شامل، بما في ذلك ضغط الأوزان على مستوى الطبقة، وتقدير النطاق القائم على المعايرة، وتعيين الذاكرة الخاص بالدقة. كما يتضمن أنظمة لتقييم أداء النموذج لقياس تأثير التكميم على الدقة عبر مهام اللغة والتلخيص.

    Implements the GPTQ algorithm for high-efficiency post-training weight quantization of large language models.

    Python
    عرض على GitHub↗5,070
  • sakurallm/sakurallmالصورة الرمزية لـ SakuraLLM

    SakuraLLM/SakuraLLM

    4,618عرض على GitHub↗

    SakuraLLM is a multi-format document translation system that hosts large language models for translating Japanese text into other languages. It functions as an inference server that exposes translation models through an OpenAI-compatible API, allowing any tool supporting the OpenAI client format to send translation requests. The system is designed as a glossary-aware translation engine that applies user-defined term dictionaries to ensure consistent translation of proper nouns and names across outputs. The project distinguishes itself by supporting multiple high-performance inference backends

    Runs the translation model on NVIDIA and AMD GPUs with CPU-GPU hybrid inference for lower-memory setups.

    Python
    عرض على GitHub↗4,618
  • baichuan-inc/baichuan2الصورة الرمزية لـ baichuan-inc

    baichuan-inc/Baichuan2

    4,098عرض على GitHub↗

    Baichuan2 is a collection of pre-trained large language models, including base and chat variants, designed for natural language generation and multi-turn conversational AI. It provides an inference engine and a fine-tuning framework to adapt these models to custom datasets and specialized domains. The project features a quantization toolkit and an inference engine that enable model execution across diverse hardware, including graphics processors, central processors, and specialized accelerators. These tools support low-bit weight quantization to reduce memory usage and increase inference spee

    Implements weight quantization to four or eight bits to reduce memory overhead and increase inference speed.

    Pythonartificial-intelligencebenchmarkceval
    عرض على GitHub↗4,098
  • nunchaku-ai/nunchakuالصورة الرمزية لـ nunchaku-ai

    nunchaku-ai/nunchaku

    3,883عرض على GitHub↗

    Nunchaku is a 4-bit model quantization library and diffusion model inference engine designed to run large-scale neural networks on consumer GPUs. It functions as a GPU-accelerated optimizer that reduces VRAM usage and increases inference speed through weight compression and memory management. The project utilizes low-rank weight decomposition and SVD weight quantization to compress models to four-bit precision while maintaining visual fidelity. It employs kernel-level operator fusion to minimize data movement and hardware-aware precision mapping to adjust numerical precision based on the unde

    Provides a toolkit for compressing neural network weights into four-bit precision to reduce VRAM usage.

    Pythoncomfyuidiffusion-modelsflux
    عرض على GitHub↗3,883
  1. Home
  2. DevOps & Infrastructure
  3. Intel Hardware Acceleration
  4. Low-Bit Weight Quantization

استكشف الوسوم الفرعية

  • 4-Bit Quantization Tools2 وسوم فرعيةTools that reduce model weight memory by approximately 70% using 4-bit quantization with minimal accuracy loss. **Distinct from Low-Bit Weight Quantization:** Distinct from Low-Bit Weight Quantization: focuses specifically on 4-bit precision rather than general low-bit formats like INT4 and FP8.
  • Consumer GPU Optimizations1 وسم فرعيSpecialized configurations to allow large models to run on consumer-grade graphics cards. **Distinct from Low-Bit Weight Quantization:** Focuses on the execution target (consumer hardware) rather than the quantization process itself.
  • Importance Matrix CalibrationTechniques that use per-column weight data from calibration text to allocate more precision to high-impact weights during quantization. **Distinct from Low-Bit Weight Quantization:** Distinct from Low-Bit Weight Quantization: focuses on the calibration method for allocating precision, not the general technique of reducing bit width.