1 Repo
Tools that recommend optimal quantization and device mapping based on model configuration and detected hardware.
Distinct from Hardware Performance Tuning: Distinct from Hardware Performance Tuning: focuses on automatic recommendation rather than manual optimization.
Explore 1 awesome GitHub repository matching artificial intelligence & ml · Automatic Hardware Tuners. Refine with filters or upvote what's useful.
mistral.rs is an inference engine for large language models that runs locally and exposes models behind OpenAI and Anthropic-compatible APIs. It serves as a multi-model serving platform, capable of loading several models in a single server process with per-request routing and on-demand loading and unloading. The engine supports multimodal inference, processing text alongside images, video, audio, and speech inputs, and includes a quantized model deployment runtime that reduces memory use and speeds up inference on consumer hardware. The project distinguishes itself through an agentic tool exe
Recommends optimal quantization and device mapping based on the model config and detected hardware.