awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

1 repositorio

Awesome GitHub RepositoriesQuantized Model Runtimes

Runtimes specifically optimized to execute models that have undergone low-bit weight quantization.

Distinct from 4-Bit Quantization Tools: Focuses on the execution/runtime phase of 4-bit models rather than the tools used to perform the quantization.

Explore 1 awesome GitHub repository matching devops & infrastructure · Quantized Model Runtimes. Refine with filters or upvote what's useful.

Awesome Quantized Model Runtimes GitHub Repositories

Encuentra los mejores repositorios con IA.Buscaremos los repositorios que mejor coincidan usando IA.
  • nunchaku-ai/nunchakuAvatar de nunchaku-ai

    nunchaku-ai/nunchaku

    3,883Ver en GitHub↗

    Nunchaku is a 4-bit model quantization library and diffusion model inference engine designed to run large-scale neural networks on consumer GPUs. It functions as a GPU-accelerated optimizer that reduces VRAM usage and increases inference speed through weight compression and memory management. The project utilizes low-rank weight decomposition and SVD weight quantization to compress models to four-bit precision while maintaining visual fidelity. It employs kernel-level operator fusion to minimize data movement and hardware-aware precision mapping to adjust numerical precision based on the unde

    Runs large-scale diffusion models on consumer-grade hardware using 4-bit quantization to maximize performance.

    Pythoncomfyuidiffusion-modelsflux
    Ver en GitHub↗3,883
  1. Home
  2. DevOps & Infrastructure
  3. Intel Hardware Acceleration
  4. Low-Bit Weight Quantization
  5. 4-Bit Quantization Tools
  6. Quantized Model Runtimes