1 repositorio
Mechanisms for simultaneously writing large model states across multiple storage nodes.
Distinct from Model Checkpointing: Focuses on the parallel streaming nature of the write operation rather than the general persistence of model weights
Explore 1 awesome GitHub repository matching artificial intelligence & ml · Parallel Checkpoint Streams. Refine with filters or upvote what's useful.
3FS is a distributed file system and RDMA storage cluster designed for high-performance AI training and inference workloads. It functions as a strongly consistent storage layer that utilizes a disaggregated architecture to pool SSDs and memory resources across multiple nodes. The system provides specialized storage implementations including an AI training checkpoint store for parallel state preservation and a distributed key-value cache store for decoder layer vectors to optimize inference processing. It ensures data integrity through chain replication and apportioned query distribution. The
Writes large model states across multiple storage nodes simultaneously to minimize training downtime during synchronization.