1 dépôt
Optimizing Hugging Face transformer models by swapping standard layers for memory-efficient Triton kernels with a single function call.
Distinct from Hugging Face: Distinct from Hugging Face: focuses on runtime layer optimization via kernel patching, not model format conversion.
Explore 1 awesome GitHub repository matching devops & infrastructure · Layer Optimization Patches. Refine with filters or upvote what's useful.
Liger-Kernel is a collection of pre-built fused Triton kernels and patching utilities designed to accelerate large language model training. It provides drop-in kernel replacements for common LLM operations such as RMSNorm, cross-entropy loss, and attention, enabling increased throughput and reduced memory usage while preserving bitwise-exact gradients. The project serves as a toolkit for composing custom model architectures from individual optimized kernels and for patching pre-existing models with minimal code changes. The project distinguishes itself through its ability to perform runtime m
Optimizes Hugging Face transformer models by swapping standard layers for memory-efficient Triton kernels with a single function call.