1 repository
Settings for managing the local cache of Hugging Face model assets, including directory paths and offline mode.
Distinct from Hugging Face: Distinct from Hugging Face: focuses on cache directory and offline configuration rather than model conversion.
Explore 1 awesome GitHub repository matching devops & infrastructure · Cache Configurations. Refine with filters or upvote what's useful.
mistral.rs is an inference engine for large language models that runs locally and exposes models behind OpenAI and Anthropic-compatible APIs. It serves as a multi-model serving platform, capable of loading several models in a single server process with per-request routing and on-demand loading and unloading. The engine supports multimodal inference, processing text alongside images, video, audio, and speech inputs, and includes a quantized model deployment runtime that reduces memory use and speeds up inference on consumer hardware. The project distinguishes itself through an agentic tool exe
Sets root and hub cache directories for Hugging Face assets with optional offline mode.