2 مستودعات
Memory management that uses a predictor model to load frequent data into high-speed memory.
Distinct from Memory Management Systems: Distinct from Memory Management Systems: specifically uses a predictor model to preload neurons into VRAM.
Explore 2 awesome GitHub repositories matching software engineering & architecture · Predictive Memory Preloading. Refine with filters or upvote what's useful.
PowerInfer is an inference engine and serving framework designed to run large language models on local hardware. It combines a hybrid CPU-GPU offloader, a quantization tool, and a sparse model optimizer to enable the execution of high-parameter models on consumer-grade devices. The system distinguishes itself through neuron-activation-based offloading, using a predictor model to preload frequent neurons into VRAM while keeping rare neurons in system memory. This hybrid execution model balances workloads between the GPU and CPU based on input patterns to optimize memory access and increase tok
Uses a small predictor model to preload frequent neurons into VRAM while keeping rare ones in system memory.
MemOS is an open-source persistent memory layer for AI agents and large language models, providing a self-hosted server that stores and retrieves structured memory across sessions. It enables AI systems to recall user preferences, history, and context without retraining, using a graph-based API and a web management interface for viewing, editing, and organizing memory items, skills, traces, and knowledge bases. The system distinguishes itself through a portable memory interchange protocol that allows memory to be transferred between different AI models, devices, and applications, along with a
Loads relevant memory before it is needed by analyzing dialogue history, task semantics, or environmental cues.