awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
google-ai-edge avatar

google-ai-edge/LiteRT-LM

0
View on GitHub↗
5,619 stars·584 forks·C++·Apache-2.0·12 viewsai.google.dev/edge/litert-lm↗

LiteRT LM

LiteRT-LM is a high-performance inference framework designed to execute large language models locally on mobile, desktop, and IoT hardware. It serves as an on-device model runtime that utilizes CPU, GPU, and NPU acceleration to provide low-latency processing.

The framework is distinguished by its ability to process text, vision, and audio inputs through a single multi-modal inference engine. It features a local HTTP server that emulates OpenAI-compatible API endpoints and a WebGPU-based runtime for executing models directly within a web browser. To ensure output reliability, it includes a constrained text generator that enforces JSON schemas or grammar rules on model responses.

The project provides broad capabilities for stateful conversation management, speculative decoding for increased token generation speeds, and a tool-calling interface that maps model requests to external functions. It also includes specialized integration for the Apple ecosystem and a dedicated plugin for running models in Flutter.

Users can execute models through a command line interface or integrate them into applications via native APIs.

Features

  • On-Device Inference - Executes large language models locally on Linux, macOS, Windows, and Raspberry Pi using local hardware.
  • Local Model Execution - Provides a high-performance framework for running large language models locally on mobile, desktop, and IoT hardware.
  • Edge AI Model Deployment - Optimizes and deploys generative AI models for high-performance execution on local CPU, GPU, and NPU hardware.
  • On-Device Model Runtimes - Provides a high-performance runtime for executing large language models locally on mobile, desktop, and IoT hardware.
  • Agent Tool Execution - Implements mechanisms for models to invoke external functions to automate complex agentic workflows.
  • AI Conversation Managers - Manages AI chat sessions including system instructions, message history, and sampling parameters.
  • Conversation Management Systems - Creates isolated conversation sessions with customizable system prompts to maintain user-specific context.
  • Conversation State Management - Tracks interaction history and context between users and models to ensure dialogue continuity.
  • Chat Model Text Generators - Generates text responses from user prompts using chat-specific configurations for local model inference.
  • External Tool Integration - Allows the model to request the execution of client-side functions to interact with external data and APIs.
  • Hardware Acceleration Backends - Implements hardware acceleration backends that route operations to CPU, GPU, and NPU for high-performance edge inference.
  • Browser-Based Inference - Executes large language models directly in the web browser using hardware-accelerated WebGPU runtimes.
  • Inference Frameworks - Provides a high-performance inference framework for executing large language models on mobile, desktop, and IoT hardware.
  • Hardware Acceleration - Leverages GPU and NPU accelerators to maximize processing speed for on-device model operations.
  • Multi-Modal Inference Engines - Provides a local runtime capable of processing text, vision, and audio inputs within a single model.
  • Multi-Modal Inference Engines - Features a single execution engine capable of processing text, vision, and audio inputs simultaneously on-device.
  • Edge Model Compilers - Converts large language models into specialized, quantized runtime files tailored for mobile, desktop, and IoT hardware.
  • Model Inference Execution - Executes data through a processing pipeline to generate linguistic outputs via blocking or streaming calls.
  • Multimodal Data Processing - Supports processing a combination of text, images, and audio data within a single inference engine.
  • On-Device Inference Engines - Provides a high-performance runtime optimized for executing large language models locally on edge hardware.
  • Grammar-Constrained Samplers - Enforces specific JSON schemas or grammar rules on model output to ensure valid response formats.
  • On-Device Inference Executions - Runs machine learning model predictions directly on local hardware accelerators across various operating systems.
  • Tool-Use Integrations - Allows defining external functions via specifications that the model can invoke to perform actions or fetch data.
  • Constrained Decoding for Tool Use - Implements constrained decoding to facilitate precise function calling and agentic workflows.
  • Schema-Constrained Sampling - Implements token selection restrictions based on JSON schemas or grammar rules during model inference.
  • Persona and Behavioral Instructions - Allows the definition of system instructions and personas to control model behavior during conversations.
  • Browser-Based Model Inference - Executes language models directly in the web browser using WebGPU for hardware-accelerated local text generation.
  • Conversation State Management - Maintains conversation state and dialogue history to provide context across multiple turns in a session.
  • State-Tracking Dialogue Managers - Tracks interaction history and applies system instructions to maintain dialogue continuity across multiple messages.
  • Tool Integrations - Maps model-generated tool requests to external client-side functions for real-time data retrieval.
  • External Tool Execution - Defines custom logic and parameters that the model can automatically invoke to perform tasks or fetch data.
  • Function Calling Interfaces - Integrates external functions into on-device models to perform specific actions or retrieve real-time data.
  • Automatic Tool Executions - Automatically invokes predefined external functions based on model requirements using docstrings and type hints.
  • Apple Silicon GPU Accelerators - Integrates language models into Apple ecosystems with specific GPU acceleration for M-series hardware.
  • In-Browser Model Execution - Executes language models directly in the browser using WebGPU for hardware-accelerated inference.
  • Prompt Templates - Uses template engines to format raw input into structured prompts compatible with instruction-tuned models.
  • Interactive Agent Chat Interfaces - Provides interfaces for real-time, conversational interaction with local AI models.
  • Browser-based Inference Engines - Executes large language models directly in the web browser using WebGPU hardware acceleration.
  • Apple Ecosystem Integration - Optimizes models for native deployment and GPU acceleration within Apple operating systems.
  • Model Capability Extensions - Provides integrations for multi-token prediction and function calling to enhance the functional output of local models.
  • Conversational Context Initialization - Defines the starting state of conversations using system instructions and initial examples.
  • Grammar-Constrained Generation - Enforces JSON schemas or grammar rules on model responses to ensure predictable and valid output formats.
  • Model Response Streaming - Streams model outputs incrementally to the client for a more responsive user experience.
  • OpenAI-Compatible Model Servers - Hosts a local HTTP server that implements the OpenAI API specification for seamless model integration.
  • On-Device Tool Integration - Connects local language models to external functions and APIs for real-time data retrieval and action execution.
  • Prompt Templates - Uses a template engine to format raw inputs into structured prompts compatible with specific models.
  • Constrained Decoding - Provides constrained decoding to ensure model outputs follow specific structured formats via logit manipulation.
  • Structured Output Enforcements - Restricts model responses to specific formats using JSON schemas, regular expressions, or grammar rules.
  • Speculative Decoding - Increases token generation speed on GPU backends by predicting multiple tokens simultaneously via speculative decoding.
  • Tool-Calling Mappings - Maps model-generated requests to external functions using descriptive docstrings and type hints for real-time data retrieval.
  • Token Mixing Accelerators - Reduces inference latency on mobile GPUs by employing multi-token prediction strategies.
  • Inference Engine API Bridges - Provides native API interfaces for integrating on-device model inference into cross-platform applications.
  • LLM Inference Plugins - Provides a dedicated plugin to run large language models locally in Flutter-based mobile and desktop applications.
  • On-Device LLM Flutter Plugins - Provides a dedicated Flutter plugin for executing large language models locally on mobile and desktop.
  • API Emulators - Hosts a local server that mimics standard API endpoints to integrate on-device inference into existing software workflows.
  • OpenAI-Compatible Servers - Implements a local server mimicking OpenAI API endpoints to ensure compatibility with existing AI software workflows.
  • OpenAI-Compatible API Servers - Provides a local HTTP server that implements the OpenAI API specification for drop-in model serving.
  • Inference Engines - Production-ready framework for deploying LLMs on edge devices.
  • Model Serving & Deployment - Deploys LLMs specifically on edge devices.

Star history

Star history chart for google-ai-edge/litert-lmStar history chart for google-ai-edge/litert-lm

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Open-source alternatives to LiteRT LM

Similar open-source projects, ranked by how many features they share with LiteRT LM.
  • pytorch/executorchpytorch avatar

    pytorch/executorch

    4,296View on GitHub↗

    ExecuTorch is a lightweight C++ runtime for deploying PyTorch models on mobile, embedded, and edge hardware. It provides an ahead-of-time compilation pipeline that exports, quantizes, and lowers model graphs into compact serialized programs, then executes them through a minimal runtime with hardware acceleration and on-device large language model inference capabilities. The project distinguishes itself through a hardware accelerator delegate system that partitions model subgraphs and offloads computation to specialized backends including NPUs, GPUs, and DSPs from Apple, Arm, Intel, MediaTek,

    Pythondeep-learningembeddedgpu
    View on GitHub↗4,296
  • microsoft/vscode-copilot-chatmicrosoft avatar

    microsoft/vscode-copilot-chat

    9,493View on GitHub↗

    This project is an AI-powered IDE extension and LLM coding assistant that provides a conversational interface for generating, refactoring, and debugging code. It functions as an AI agent framework and a Model Context Protocol client, connecting AI models to external data sources and tools to automate complex development tasks. The system is distinguished by its use of autonomous AI agents capable of multi-step task execution, including the ability to read files, modify code, and run terminal commands iteratively. It supports recursive agent orchestration through subagent delegation and employ

    TypeScript
    View on GitHub↗9,493
  • vercel/aivercel avatar

    vercel/ai

    21,885View on GitHub↗

    This project is a comprehensive framework for building AI-powered applications, providing a unified toolkit for orchestrating language models, autonomous agents, and interactive user interfaces. It serves as a central library for managing the entire lifecycle of AI interactions, from initial prompt generation and model provider abstraction to complex, multi-step reasoning and tool execution. The framework distinguishes itself through its deep integration with frontend development, specifically by enabling generative user interfaces that render dynamic components directly from model outputs. I

    TypeScriptanthropicartificial-intelligencegemini
    View on GitHub↗21,885
  • agiresearch/aiosagiresearch avatar

    agiresearch/AIOS

    5,168View on GitHub↗

    AIOS is an LLM agent operating system and orchestration kernel designed to manage memory, resource scheduling, and tool execution for multiple autonomous AI agents. It serves as a comprehensive framework for developing and deploying agents, featuring a dedicated resource manager that coordinates model backends, GPU memory, and isolated kernel instances. The system distinguishes itself through a semantic memory engine that uses vector search and autonomous clustering for long-term knowledge management, and a semantic file system that allows users to control computer files and system operations

    Python
    View on GitHub↗5,168
See all 30 alternatives to LiteRT LM→

Frequently asked questions

What does google-ai-edge/litert-lm do?

LiteRT-LM is a high-performance inference framework designed to execute large language models locally on mobile, desktop, and IoT hardware. It serves as an on-device model runtime that utilizes CPU, GPU, and NPU acceleration to provide low-latency processing.

What are the main features of google-ai-edge/litert-lm?

The main features of google-ai-edge/litert-lm are: On-Device Inference, Local Model Execution, Edge AI Model Deployment, On-Device Model Runtimes, Agent Tool Execution, AI Conversation Managers, Conversation Management Systems, Conversation State Management.

What are some open-source alternatives to google-ai-edge/litert-lm?

Open-source alternatives to google-ai-edge/litert-lm include: pytorch/executorch — ExecuTorch is a lightweight C++ runtime for deploying PyTorch models on mobile, embedded, and edge hardware. It… microsoft/vscode-copilot-chat — This project is an AI-powered IDE extension and LLM coding assistant that provides a conversational interface for… vercel/ai — This project is a comprehensive framework for building AI-powered applications, providing a unified toolkit for… agiresearch/aios — AIOS is an LLM agent operating system and orchestration kernel designed to manage memory, resource scheduling, and… microsoft/onnxruntime — This project is a cross-platform machine learning inference engine designed to execute pre-trained models across… openai/openai-agents-python — This project is a Python framework for building autonomous, event-driven agent systems. It provides a unified runtime…