1 repository
8-bit model quantization using the bitsandbytes library for reduced memory inference with explicit CUDA tensor management.
Distinct from Quantized Model Implementations: Distinct from Quantized Model Implementations: focuses on the specific bitsandbytes library and its 8-bit quantization technique, not general quantized model versions.
Explore 1 awesome GitHub repository matching artificial intelligence & ml · BitsAndBytes Quantizers. Refine with filters or upvote what's useful.
Provides 8-bit quantization via bitsandbytes for memory-efficient inference on CUDA devices.