1 dépôt
Memory-reduction techniques using low-rank approximations of momentum matrices, such as AdaFactor.
Distinct from Training Memory Optimizers: Focuses specifically on the optimizer's internal state compression for memory efficiency, distinct from general gradient checkpointing.
Explore 1 awesome GitHub repository matching data & databases · Optimizer State Compression. Refine with filters or upvote what's useful.
This project is a comprehensive educational curriculum and structured learning path covering the full lifecycle of large language models. It provides a guided progression through the theory, architecture, training, and deployment of these models. The curriculum includes specialized guides on transformer architecture, model training tutorials, and frameworks for designing autonomous agents. It also provides dedicated resources for studying model safety and ethics. The material covers a wide range of technical capabilities, including distributed training strategies, parameter-efficient fine-tu
Details the use of AdaFactor to reduce training memory via low-rank momentum approximations.