awesome-repositories.com
Blog
MCP
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetServeur MCPÀ proposNotre méthodologiePresse
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
RunanywhereAI avatar

RunanywhereAI/runanywhere-sdks

0
View on GitHub↗
8,781 stars·265 forks·C++·other·15 vueswww.runanywhere.ai↗

Runanywhere Sdks

This project is an on-device AI SDK providing a framework for running large language models, vision models, and speech models locally. It serves as an orchestration layer for local LLM execution, ensuring data privacy and offline availability by utilizing hardware acceleration on the device.

The SDK is distinguished by its comprehensive voice and multimodal capabilities, including a coordinated voice pipeline for activity detection, speech-to-text, and text-to-speech synthesis. It also provides a dedicated implementation kit for local retrieval-augmented generation and tools for processing combined image and text inputs via vision-language models.

The broader capability surface covers model lifecycle management, including downloading, caching, and the dynamic swapping of fine-tuned adapters. It includes support for structured output generation, tool calling for external function integration, and hardware-accelerated image generation.

The system also incorporates performance monitoring for inference metrics and comprehensive audio-visual capture tools for camera and microphone input.

Features

  • Local Model Execution - Enables the execution of large language and vision models locally on hardware for privacy and offline availability.
  • On-Device Inference Engines - Implements a high-performance runtime for executing large language and vision models locally using hardware acceleration.
  • Voice Pipelines - Implements a coordinated offline pipeline for voice activity detection, speech-to-text, and text-to-speech synthesis.
  • Voice Interaction Management - Maintains a continuous voice session with automatic activity detection and event-based callbacks.
  • AI Integration Tools - Connects local models to external functions and custom code via structured tool calling and argument parsing.
  • Speech-to-Text Translation - Converts mono PCM audio into text using specialized speech-to-text models.
  • Conversational Voice AI - Provides a coordinated pipeline for building hands-free assistants using voice activity detection, speech-to-text, and text-to-speech.
  • Conversational Voice Pipelines - Coordinates voice activity detection, speech-to-text, and text-to-speech into a continuous conversational loop.
  • GPU Acceleration - Utilizes hardware-specific GPU drivers to increase the processing speed of local AI models.
  • Hardware-Accelerated Inference - Provides mechanisms to verify if the inference engine is utilizing hardware-accelerated GPU execution.
  • Local Model Orchestrators - Orchestrates the lifecycle, hardware acceleration, and streaming execution of local machine learning models.
  • Local RAG Implementations - Implements retrieval-augmented generation using private local datasets and offline LLMs for grounded answers.
  • Local Model Lifecycle Managers - Provides tools for downloading, caching, and loading model assets to optimize local device storage and RAM.
  • Multimodal Vision Inputs - Provides tools to process and interpret combined image and text inputs for visual analysis and descriptions.
  • Natural Language Generation - Produces human-like natural language responses from prompts using configurable temperature and token limits.
  • On-Device Models - Ships a comprehensive SDK for running large language, vision, and speech models locally on edge hardware.
  • Real-Time Audio Transcribers - Transcribes live audio input and emits partial results for immediate user feedback.
  • Incremental Synthesis - Produces audio chunks incrementally so playback begins before the entire synthesis task completes.
  • Local Speech Synthesis - Converts text into natural-sounding audio using neural models running locally on-device.
  • Tool Calling - Connects on-device models to external functions via typed tool definitions and automated execution.
  • Vector Embeddings - Provides vector embeddings of text to enable semantic search and memory operations directly on-device.
  • Vector Retrieval Systems - Implements local document chunking and similarity search via embedding models to provide grounded AI responses.
  • Vision-Language Inference - Executes multimodal models that process combined image and text inputs to generate analytical descriptions.
  • Voice Activity Detection - Identifies when a user starts and stops speaking using configurable sensitivity thresholds.
  • Voice Conversational Loops - Coordinates a single cycle of audio-to-text, AI reasoning, and text-to-audio synthesis.
  • Local Model Loading - Handles the downloading of model files from remote URLs and loading them into device memory.
  • Multimodal Analysis Engines - Implements a multimodal analysis engine to process visual content and text prompts together for image description.
  • Word-Level Timestamps - Tracks start and end times for individual words to synchronize text and audio playback.
  • Retrieval-Augmented Generation - Implements a retrieval-augmented generation pipeline that grounds AI responses using locally indexed documents.
  • Argument Parsing - Extracts tool names and arguments from raw model text for manual execution logic.
  • Diffusion Model Managers - Manages the registration and downloading of diffusion model packages for local image generation.
  • Inference Telemetry - Tracks tokens used and generation speed to monitor local model performance and efficiency.
  • Image Generation - Creates images from text prompts using hardware-accelerated diffusion models on-device.
  • Model Preloading Endpoints - Loads specific models into memory during idle time to reduce latency when starting a user task.
  • Incremental Inference Streaming - Streams model outputs token-by-token using asynchronous iterators to reduce perceived user latency.
  • Model Adapters - Injects lightweight LoRA adapters into base models at runtime to change behavior without reloading full weights.
  • Behavioral Constraints - Defines personas, operational constraints, and response formats for models through system prompts.
  • Model Persistence - Caches downloaded models in persistent storage to prevent redundant network requests.
  • LoRA Adapter Loaders - Applies lightweight LoRA adapters to base models at runtime to modify behavior without full reloading.
  • Structured Output Generators - Constrains model responses to specific JSON schemas to ensure machine-readable and type-safe outputs.
  • Speaker Diarization - Detects and labels different speakers within a single recording to distinguish between voices.
  • Concurrent Model Execution - Supports the simultaneous execution of multiple models in memory to facilitate complex processing pipelines.
  • Text Generation Controls - Provides controls for adjusting creativity, token limits, and streaming behavior during text generation.
  • Model Memory Reclamation - Removes unused large language or speech models from memory to reclaim system resources.
  • Document Ingestion Pipelines - Provides a pipeline for chunking, embedding, and indexing raw documents to facilitate local vector search.
  • Schema-Constrained Sampling - Enforces structured JSON or XML output formats by constraining token sampling during the model generation process.
  • Raw Audio Captures - Records raw audio at specified sample rates, providing audio chunks and volume levels for AI processing.
  • Camera Feed Capture - Accesses the device camera to provide live RGB frames for analysis by on-device vision models.
  • Token Streaming - Delivers AI model generated tokens to the user interface in real-time via asynchronous streams.
  • Event Monitoring Streams - Provides an event stream to track AI-specific lifecycle events, including generation status and model loading.
  • Inference Performance Monitoring - Tracks inference-specific metrics such as tokens per second, latency, and time to first token.
  • Model Serving & Deployment - Runs AI models on-device for mobile platforms.

Historique des stars

Graphique de l'historique des stars pour runanywhereai/runanywhere-sdksGraphique de l'historique des stars pour runanywhereai/runanywhere-sdks

Recherche par IA

Explorez plus de dépôts awesome

Décrivez vos besoins en langage naturel — l'IA classe des milliers de projets open source sélectionnés par pertinence.

Start searching with AI

Alternatives open source à Runanywhere Sdks

Projets open source similaires, classés selon le nombre de fonctionnalités partagées avec Runanywhere Sdks.
  • cactus-compute/cactusAvatar de cactus-compute

    cactus-compute/cactus

    5,363Voir sur GitHub↗

    Cactus is an on-device AI inference engine designed for executing large language models, vision models, and speech-to-text systems on mobile and wearable hardware. It provides a programmable tensor computation graph for defining sequences of matrix operations and activation functions, alongside a local retrieval augmented generation framework that grounds model responses using local text files. The project features a multiplatform SDK with language bindings for integrating AI capabilities into mobile applications and a model conversion system that transforms external model formats for optimiz

    C++aiandroidarm
    Voir sur GitHub↗5,363
  • vocodedev/vocode-coreAvatar de vocodedev

    vocodedev/vocode-core

    3,693Voir sur GitHub↗

    Vocode-core is a framework for building real-time conversational AI voice agents. It serves as a conversational orchestrator and pipeline that integrates speech-to-text, large language models, and text-to-speech services to enable low-latency voice interactions. The project features a provider-agnostic interface that allows for swappable speech and language model providers, including support for both cloud APIs and local binaries. It distinguishes itself through a specialized telephony integration layer that enables agents to be deployed across phone lines, WebRTC, and virtual meeting platfor

    Python
    Voir sur GitHub↗3,693
  • crmne/ruby_llmAvatar de crmne

    crmne/ruby_llm

    3,566Voir sur GitHub↗

    ruby_llm is an LLM integration framework and AI agent orchestrator designed to connect applications to multiple large language model providers through a unified interface. It serves as a toolkit for building autonomous assistants with custom personas, managing structured output via JSON schemas, and implementing vector embedding engines for semantic search. The project distinguishes itself as an observability suite and multimodal toolkit. It provides specialized capabilities for tracking token usage, calculating model costs, and tracing workflows via OpenTelemetry, while supporting the proces

    Rubyaianthropicchatgpt
    Voir sur GitHub↗3,566
  • nexaai/nexa-sdkAvatar de NexaAI

    NexaAI/nexa-sdk

    7,721Voir sur GitHub↗

    The nexa-sdk is an on-device AI SDK and multimodal inference engine designed to run large language, vision, and audio models locally on mobile and desktop hardware. It functions as a local LLM runtime and NPU acceleration framework, enabling the execution of generative and discriminative models without reliance on cloud services. The project distinguishes itself through a dedicated NPU acceleration framework that optimizes model execution on Neural Processing Units to reduce latency and power consumption. It employs hardware-agnostic backend routing to dynamically distribute computations acro

    Kotlingemma3gogpt-oss
    Voir sur GitHub↗7,721
Voir les 30 alternatives à Runanywhere Sdks→

Questions fréquentes

Que fait runanywhereai/runanywhere-sdks ?

This project is an on-device AI SDK providing a framework for running large language models, vision models, and speech models locally. It serves as an orchestration layer for local LLM execution, ensuring data privacy and offline availability by utilizing hardware acceleration on the device.

Quelles sont les fonctionnalités principales de runanywhereai/runanywhere-sdks ?

Les fonctionnalités principales de runanywhereai/runanywhere-sdks sont : Local Model Execution, On-Device Inference Engines, Voice Pipelines, Voice Interaction Management, AI Integration Tools, Speech-to-Text Translation, Conversational Voice AI, Conversational Voice Pipelines.

Quelles sont les alternatives open-source à runanywhereai/runanywhere-sdks ?

Les alternatives open-source à runanywhereai/runanywhere-sdks incluent : cactus-compute/cactus — Cactus is an on-device AI inference engine designed for executing large language models, vision models, and… vocodedev/vocode-core — Vocode-core is a framework for building real-time conversational AI voice agents. It serves as a conversational… crmne/ruby_llm — ruby_llm is an LLM integration framework and AI agent orchestrator designed to connect applications to multiple large… nexaai/nexa-sdk — The nexa-sdk is an on-device AI SDK and multimodal inference engine designed to run large language, vision, and audio… dusty-nv/jetson-inference — jetson-inference is a set of libraries and tools for executing optimized deep learning models on embedded GPU… aaswordman/operit — Operit is a private, voice-enabled AI agent designed to run quantized large language models offline within mobile…