awesome-repositories.com
Blog
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoAcerca deCómo clasificamosPrensaServidor MCP
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
ggml-org avatar

ggml-org/llama.cpp

0
View on GitHub↗
116,799 estrellas·19,628 forks·C++·MIT·27 vistas

Llama.cpp

Llama.cpp is an inference engine designed for the local execution of text-based and multimodal language models on consumer hardware. It provides a core environment for running models that process both text and image inputs, utilizing hardware-accelerated backends to optimize performance across diverse CPU and GPU architectures.

The project distinguishes itself by offering a lightweight HTTP server that adheres to standard API specifications, enabling chat completion, embeddings, and reranking services. It includes a suite of tools for model quantization and conversion, which reduces memory usage and improves performance, alongside a command-line interface for managing chat templates and inference parameters.

The ecosystem further supports structured data generation through grammar-based output constraints and provides diagnostic utilities for visualizing computational graphs. Comprehensive documentation is available, including a reference matrix that details the compatibility of computational operations across supported hardware backends.

Features

  • Text-Only Inference Engines - Executes large language models locally on standard consumer hardware with high performance.
  • Hardware Abstraction Layers - Unifies diverse CPU and GPU architectures through a common interface to normalize model execution across heterogeneous hardware.
  • Multimodal Inference Engines - Processes both text and image inputs locally to enable multimodal model capabilities on standard consumer devices.
  • Inference API Servers - Exposes inference capabilities via a lightweight HTTP server that supports standard chat completion and embedding endpoints.
  • Model Quantization Tools - Compresses model weights into quantized formats to significantly reduce memory footprint and boost inference speed.
  • AI and Machine Learning - Efficient inference engine for large language models.
  • AI & Machine Learning - LLM inference in C/C++.
  • Inference and Serving - High-performance inference engine written in C/C++.
  • Inference Engines - Efficient LLM inference implementation in C/C++.
  • Large Language Models - High-performance LLM inference in C/C++.
  • Local Development and Serving - Execute high-performance model inference across various hardware backends.
  • Model Quantization - Listed in the “Model Quantization” section of the Llm Course awesome list.
  • Model Serving & Deployment - Performs efficient local inference for various LLMs.
  • Running Models - Listed in the “Running Models” section of the Llm Course awesome list.
  • Command Line Inference Interfaces - Terminal-based utilities allow for direct interaction with models, including configuration of inference parameters and chat management.

Historial de estrellas

Gráfico del historial de estrellas de ggml-org/llama.cppGráfico del historial de estrellas de ggml-org/llama.cpp

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Start searching with AI

Alternativas open-source a Llama.cpp

Proyectos open-source similares, clasificados según cuántas características comparten con Llama.cpp.
  • berriai/litellmAvatar de BerriAI

    BerriAI/litellm

    50,579Ver en GitHub↗

    LiteLLM is a unified gateway and proxy server designed to centralize access to over one hundred language model providers. It provides a standardized API interface that abstracts vendor-specific schemas, allowing developers to interact with diverse models through a single, consistent format. By acting as a central traffic management layer, it enables organizations to route, secure, and govern model interactions across multiple deployments. The platform distinguishes itself through its policy-driven architecture, which uses configuration-based routing to manage traffic distribution, load balanc

    Pythonai-gatewayanthropicazure-openai
    Ver en GitHub↗50,579
  • ggerganov/llama.cppAvatar de ggerganov

    ggerganov/llama.cpp

    116,912Ver en GitHub↗

    llama.cpp is a high-performance C++ inference engine and runtime for executing large language models locally across various hardware architectures. It provides the core components for local model execution, including a dedicated model quantizer for compressing weights into the GGUF format and a system for generating text embeddings for semantic search. The project distinguishes itself through specialized memory and execution optimizations, such as block-wise weight quantization to reduce memory footprints and memory-mapped model loading. It supports structured text generation by using formal

    C++
    Ver en GitHub↗116,912
  • mudler/localaiAvatar de mudler

    mudler/LocalAI

    46,889Ver en GitHub↗

    LocalAI is a self-hosted inference server that enables the execution of machine learning models directly on local hardware. By providing a unified interface for text, image, and audio processing, it allows users to maintain full control over data privacy and infrastructure costs while eliminating dependencies on external network services. The platform functions as an API gateway that mimics standard cloud-based artificial intelligence interfaces, allowing existing applications to integrate local models as drop-in replacements. It utilizes a container-based architecture to package runtimes and

    Goaiapiaudio-generation
    Ver en GitHub↗46,889
  • ollama/ollamaAvatar de ollama

    ollama/ollama

    174,300Ver en GitHub↗

    Ollama provides a framework for running and managing local machine learning models. It includes a command-line interface for model lifecycle management, such as creation, embedding generation, and configuration, alongside a stable API for programmatic interaction across multiple programming languages. The platform supports the import of models and adapters in various formats, including GGUF and Safetensors. Users can define custom model behaviors, prompt templates, and system messages through a configuration file format. It also offers tools for fine-tuning models with LoRA adapters and apply

    Godeepseekgemmagemma3
    Ver en GitHub↗174,300
Ver las 30 alternativas a Llama.cpp→

Preguntas frecuentes

¿Qué hace ggml-org/llama.cpp?

Llama.cpp is an inference engine designed for the local execution of text-based and multimodal language models on consumer hardware. It provides a core environment for running models that process both text and image inputs, utilizing hardware-accelerated backends to optimize performance across diverse CPU and GPU architectures.

¿Cuáles son las características principales de ggml-org/llama.cpp?

Las características principales de ggml-org/llama.cpp son: Text-Only Inference Engines, Hardware Abstraction Layers, Multimodal Inference Engines, Inference API Servers, Model Quantization Tools, AI and Machine Learning, AI & Machine Learning, Inference and Serving.

¿Qué alternativas de código abierto existen para ggml-org/llama.cpp?

Las alternativas de código abierto para ggml-org/llama.cpp incluyen: berriai/litellm — LiteLLM is a unified gateway and proxy server designed to centralize access to over one hundred language model… ollama/ollama — Ollama provides a framework for running and managing local machine learning models. It includes a command-line… ggerganov/llama.cpp — llama.cpp is a high-performance C++ inference engine and runtime for executing large language models locally across… mudler/localai — LocalAI is a self-hosted inference server that enables the execution of machine learning models directly on local… sgl-project/sglang — Sglang is a high-performance inference engine and serving system designed for large language and multimodal models. It… lyogavin/airllm — Airllm is a framework designed to execute and fine-tune large language models on consumer-grade hardware. By employing…