3 repository-uri
Frameworks that reduce the memory footprint of diffusion models for inference on consumer hardware.
Distinct from Diffusion Weight Optimizers: Focuses on inference memory optimization and precision recovery rather than training weight optimization.
Explore 3 awesome GitHub repositories matching artificial intelligence & ml · Diffusion Model Memory Optimizers. Refine with filters or upvote what's useful.
Text2Video-Zero este un model de difuzie text-to-video și un framework conceput pentru a sintetiza secvențe video temporal consistente din prompt-uri textuale. Funcționează ca un generator video zero-shot, reutilizând modelele de difuzie de imagine pre-antrenate pentru a crea conținut video fără a necesita antrenament suplimentar pe seturi de date video. Sistemul include un sintetizator video condițional care permite generarea ghidată folosind hărți de adâncime, margini sau postură pentru a controla layout-ul structural și mișcarea. De asemenea, oferă capabilități de editare video bazate pe text pentru a modifica stilul sau conținutul clipurilor video existente prin instrucțiuni în limbaj natural. Pentru a gestiona cerințele computaționale, proiectul implementează inferența optimizată pentru memoria GPU. Acest lucru este realizat prin tehnici precum token merging și frame chunking pentru a reduce utilizarea VRAM în timpul procesului de generare.
Optimizes GPU memory usage during video generation to enable inference on consumer-grade hardware.
Nunchaku is a 4-bit model quantization library and diffusion model inference engine designed to run large-scale neural networks on consumer GPUs. It functions as a GPU-accelerated optimizer that reduces VRAM usage and increases inference speed through weight compression and memory management. The project utilizes low-rank weight decomposition and SVD weight quantization to compress models to four-bit precision while maintaining visual fidelity. It employs kernel-level operator fusion to minimize data movement and hardware-aware precision mapping to adjust numerical precision based on the unde
Provides a specialized inference engine for running large-scale diffusion models with reduced memory overhead on consumer GPUs.
ComfyUI-GGUF is a memory optimizer and model loader for ComfyUI that enables the execution of large transformer-based generative models using quantized weights. It provides a system for loading GGUF formatted weights within a node-based diffusion interface to reduce GPU memory consumption. The project includes a quantization tool for converting standard model checkpoints into compressed binary formats and a tensor fixer to restore missing keys and correct architectures in binary model files. These utilities ensure that compressed models remain functional during inference on hardware with limi
Provides a framework for running large transformer-based generative models using quantized weights.