1 repository
Lowers peak memory usage from quadratic to linear in sequence length by processing attention in tiled chunks.
Distinct from Long-Context Sequence Processors: Distinct from Long-Context Sequence Processors: focuses on the tiled attention computation technique for memory reduction rather than general sequence processing.
Explore 1 awesome GitHub repository matching data & databases · Tiled Attention Memory Reducers. Refine with filters or upvote what's useful.
Lowers peak memory usage from quadratic to linear by processing attention in tiled chunks.