1 रिपॉजिटरी
Allocates and manages GPU buffers with unicast and multicast access for multi-node communication, exposing PyTorch tensor views.
Distinct from Multi-Node: Distinct from Multi-Node: focuses on buffer allocation and management with multicast access, not just NCCL communicator setup.
Explore 1 awesome GitHub repository matching artificial intelligence & ml · GPU Buffer Allocations. Refine with filters or upvote what's useful.
FlashInfer is a library of high-performance GPU kernels purpose-built for accelerating large language model inference. It provides optimized implementations for attention operations (including flash attention, page attention, multi-head latent attention, and cascade attention) using paged key-value caches, fused kernel composition, and just-in-time compilation. The library also includes specialized kernels for mixture-of-experts layers, block-scaled low-precision quantization (FP8, FP4), and distributed collective communication. What distinguishes FlashInfer is its fused all-reduce communicat
Provides GPU buffer allocation with multicast access for multi-node communication in distributed inference.