awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
cactus-compute avatar

cactus-compute/cactus

0
View on GitHub↗
5,363 stars·431 forks·C++·35 viewscactuscompute.com↗

Cactus

Cactus is an on-device AI inference engine designed for executing large language models, vision models, and speech-to-text systems on mobile and wearable hardware. It provides a programmable tensor computation graph for defining sequences of matrix operations and activation functions, alongside a local retrieval augmented generation framework that grounds model responses using local text files.

The project features a multiplatform SDK with language bindings for integrating AI capabilities into mobile applications and a model conversion system that transforms external model formats for optimized local execution. It utilizes a hybrid routing system to redirect workloads between on-device execution and cloud-based providers based on hardware capacity.

The engine covers a broad capability surface including on-device audio processing for voice activity detection and transcription, vector embedding generation for similarity search, and tool integration for parsing model outputs into external function calls. These processes are supported by optimized native kernels tuned for low-latency performance on mobile hardware.

Features

  • Local AI Inference - Executes large language and vision models directly on mobile and wearable hardware using optimized kernels.
  • On-Device Inference Engines - Serves as an on-device AI inference engine for executing large language, vision, and speech models on mobile and wearable hardware.
  • Multimodal Input Processing - Performs inference on image and sound data to enable visual understanding and speech-to-text capabilities.
  • Chat Completion Services - Produces natural language conversational responses based on chat history and configurable generation options.
  • Retrieval-Augmented Generation - Grounds model responses using locally stored text documents and directories to provide context-aware generation.
  • RAG Document Retrieval - Retrieves relevant snippets from local text files to provide grounded context for LLM responses.
  • Inference Optimization Kernels - Utilizes native kernels tuned for low-latency, energy-efficient mathematical operations on mobile hardware.
  • Local RAG Implementations - Provides a local retrieval augmented generation framework that grounds model responses using local text files without cloud access.
  • Local Inference Engines - Provides an optimized runtime for executing large language models and vision models locally on consumer mobile hardware.
  • On-Device Speech-to-Text SDKs - Provides on-device speech-to-text transcription using locally executed models on mobile and wearable hardware.
  • RAG Frameworks - Provides a framework for building local retrieval augmented generation systems that ground responses in local directories.
  • Speech Transcription - Provides local on-device speech-to-text transcription services with low-latency execution.
  • Vector Embeddings - Generates numerical vector representations of text, visual, and speech inputs for similarity search and retrieval.
  • AI Integration SDKs - Ships a multiplatform SDK with language bindings for integrating local AI capabilities into mobile applications.
  • Mobile Framework Integrations - Offers native software kits to integrate AI capabilities into handheld and wearable operating systems.
  • Tensor Computation Graphs - Allows defining sequences of tensor operations and activation functions as computational graphs for local execution.
  • Graph-Based Execution Engines - Executes mathematical workflows as a sequence of tensor operations and activation functions via directed acyclic graphs.
  • Language Bindings - Provides multiplatform software development kits and language bindings to connect the core engine to external applications.
  • AI Integration Tools - Connects local AI models to external system functions and tools to perform actions based on model outputs.
  • Model Request Routing - Redirects inference requests to cloud providers when local hardware capacity is insufficient.
  • Cross-Framework Model Conversion - Transforms external model formats into representations optimized for mobile and wearable hardware.
  • Function Calling Interfaces - Parses model outputs into structured function calls to interact with external system tools.
  • Hybrid Local-Remote AI Routing - Routes AI workloads between local on-device execution and cloud-based providers based on hardware capacity.
  • Local Speech-to-Text - Includes a low-latency on-device transcription system for converting audio input into text.
  • On-Device Speech Recognizers - Performs local speech-to-text transcription and voice activity detection on handheld and wearable devices.
  • Voice Activity Detection - Identifies periods of human speech within audio streams to trigger transcription and downstream processing.
  • Mobile Model Format Converters - Transforms external model formats into optimized representations compatible with local mobile and wearable hardware execution.

Star history

Star history chart for cactus-compute/cactusStar history chart for cactus-compute/cactus

How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Open-source alternatives to Cactus

Similar open-source projects, ranked by how many features they share with Cactus.
  • runanywhereai/runanywhere-sdksRunanywhereAI avatar

    RunanywhereAI/runanywhere-sdks

    8,781View on GitHub↗

    This project is an on-device AI SDK providing a framework for running large language models, vision models, and speech models locally. It serves as an orchestration layer for local LLM execution, ensuring data privacy and offline availability by utilizing hardware acceleration on the device. The SDK is distinguished by its comprehensive voice and multimodal capabilities, including a coordinated voice pipeline for activity detection, speech-to-text, and text-to-speech synthesis. It also provides a dedicated implementation kit for local retrieval-augmented generation and tools for processing co

    C++androidapple-intelligencecpp
    View on GitHub↗8,781
  • langroid/langroidlangroid avatar

    langroid/langroid

    3,894View on GitHub↗

    Langroid is a multi-agent orchestration framework and tool integration suite designed for building complex AI applications. It serves as a multi-modal integration layer that connects diverse local and remote language models with an agentic retrieval-augmented generation system. The project distinguishes itself through a collaborative message-exchange paradigm, allowing specialized agents to delegate tasks hierarchically and coordinate via structured communication. It features an advanced state management system for conversational AI, including the ability to rewind and prune conversation hist

    Pythonagentsaichatgpt
    View on GitHub↗3,894
  • imclumsypanda/langchain-chatglmimClumsyPanda avatar

    imClumsyPanda/langchain-ChatGLM

    38,183View on GitHub↗

    This project is a LangChain-based framework for building retrieval-augmented generation systems, autonomous agents, and multimodal chatbots. It functions as an open-source orchestrator that connects local inference engines and online APIs to manage various large language model deployments. The system distinguishes itself by providing specialized interfaces for local knowledge bases, allowing the loading and vectorization of private documents to create context-aware assistants. It also supports multimodal capabilities, enabling the processing of both text and image inputs through vision-capabl

    Python
    View on GitHub↗38,183
  • pipecat-ai/pipecatpipecat-ai avatar

    pipecat-ai/pipecat

    12,846View on GitHub↗

    Pipecat is a framework and software development kit for building real-time multimodal AI agents and speech-to-speech systems. It utilizes a frame-based data pipeline to route audio, video, and text through a modular sequence of processors, enabling the orchestration of low-latency conversational AI. The project is distinguished by its ability to coordinate complex multimodal services, including speech-to-text, language models, and text-to-speech, within a single pipeline. It features semantic voice activity detection for natural turn-taking, state-machine conversation flows for dialogue manag

    Pythonaichatbot-frameworkchatbots
    View on GitHub↗12,846
See all 30 alternatives to Cactus→

Frequently asked questions

What does cactus-compute/cactus do?

Cactus is an on-device AI inference engine designed for executing large language models, vision models, and speech-to-text systems on mobile and wearable hardware. It provides a programmable tensor computation graph for defining sequences of matrix operations and activation functions, alongside a local retrieval augmented generation framework that grounds model responses using local text files.

What are the main features of cactus-compute/cactus?

The main features of cactus-compute/cactus are: Local AI Inference, On-Device Inference Engines, Multimodal Input Processing, Chat Completion Services, Retrieval-Augmented Generation, RAG Document Retrieval, Inference Optimization Kernels, Local RAG Implementations.

What are some open-source alternatives to cactus-compute/cactus?

Open-source alternatives to cactus-compute/cactus include: runanywhereai/runanywhere-sdks — This project is an on-device AI SDK providing a framework for running large language models, vision models, and speech… langroid/langroid — Langroid is a multi-agent orchestration framework and tool integration suite designed for building complex AI… imclumsypanda/langchain-chatglm — This project is a LangChain-based framework for building retrieval-augmented generation systems, autonomous agents,… pipecat-ai/pipecat — Pipecat is a framework and software development kit for building real-time multimodal AI agents and speech-to-speech… xusenlinzy/api-for-open-llm — This project provides a unified server environment and gateway for hosting and executing open-source large language… timescale/pgai — pgai is a PostgreSQL AI toolkit and framework designed to integrate large language models and vector embeddings…

Curated searches featuring Cactus

Hand-picked collections where Cactus appears.
  • edge computing platform