1 Repo
Offloading mechanisms that move KV cache blocks to a shared filesystem to decouple capacity from local GPU memory.
Distinct from Activation and KV Cache Offloaders: Focuses on shared network storage rather than local DRAM or NVMe offloading
Explore 1 awesome GitHub repository matching artificial intelligence & ml · Shared Storage Offloading. Refine with filters or upvote what's useful.
llm-d is a distributed serving framework designed for large language model inference. It functions as an inference orchestrator and gateway, providing a control plane for deploying model replicas and managing hardware accelerators. The system includes a batch inference scheduler and a cache manager to coordinate request flow and memory utilization. The project is distinguished by a disaggregated serving architecture that separates prefill and decode execution phases across specialized workers to maximize throughput. It employs a hardware-agnostic control plane and tiered cache offloading, mov
Moves memory blocks to a shared file system to decouple cache capacity from local GPU memory.