How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.
DeepSpeed is a distributed deep learning optimization library and framework designed for the training and inference of massive AI models. It serves as a model parallelism orchestrator and a toolkit for scaling large language models across multiple GPUs and compute nodes. The project distinguishes itself through 3D parallelism orchestration, which combines data, pipeline, and tensor parallelism. It utilizes ZeRO-based memory partitioning to eliminate redundant storage and employs CPU-offload memory management to move weights and optimizer states to system RAM. Additionally, it provides special
MLSys 2024 Best Paper Award AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Code for the ICLR 2023 paper "GPTQ: Accurate Post-training Quantization of Generative Pretrained Transformers".
ICML 2023 SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
The official implementation of the EMNLP 2023 paper LLM-FP4
The main features of nbasyl/llm-fp4 are: Model Quantization Tools, Quantization Frameworks, Model Compression.
Open-source alternatives to nbasyl/llm-fp4 include: mit-han-lab/llm-awq — [MLSys 2024 Best Paper Award] AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration. nbasyl/ofq. ist-daslab/gptq — Code for the ICLR 2023 paper "GPTQ: Accurate Post-training Quantization of Generative Pretrained Transformers". microsoft/deepspeed — DeepSpeed is a distributed deep learning optimization library and framework designed for the training and inference of… mit-han-lab/smoothquant — [ICML 2023] SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models. snu-mllab/guidedquant — Official PyTorch implementation of "GuidedQuant: Large Language Model Quantization via Exploiting End Loss Guidance"…